Subramanian Ramanathan

dblp:10/6265 · also Ramanathan Subramanian · DBLP profile ↗
← Back
65ranked-venue papers
13as first author
13since 2021 · last 2025
0000-0001-9441-7074ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 4 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
abstract
We present EditIQ, a completely automated framework for cinematically editing scenes captured via a stationary, large field-of-view and high-resolution camera. From the static camera feed, EditIQ initially generates multiple virtual feeds, emulating a team of cameramen. These virtual camera shots termed rushes are subsequently assembled using an automated editing algorithm, whose objective is to present the viewer with the most vivid scene content. To understand key scene elements and guide the editing process, we employ a two-pronged approach: (1) a large language model (LLM)-based dialogue understanding module to analyze conversational flow, coupled with (2) visual saliency prediction to identify meaningful scene elements and camera shots therefrom. We then formulate cinematic video editing as an energy minimization problem over shot selection, where cinematic constraints determine shot choices, transitions, and continuity. EditIQ synthesizes an aesthetically and visually compelling representation of the original narrative while maintaining cinematic coherence and a smooth viewing experience. Efficacy of EditIQ against competing baselines is demonstrated via a psychophysical study involving twenty participants on the BBC Old School dataset plus eleven theatre performance videos. Video samples from EditIQ can be found at https://editiq-ave.github.io/.
Rohit Girmaji, Bhav Beri, Subramanian Ramanathan, Vineet Gandhi
IUI3
2025 MIP-GAF: A MLLM-Annotated Benchmark for Most Important Person Localization and Group Context Understanding
abstract
Estimating the Most Important Person (MIP) in any social event setup is a challenging problem mainly due to contextual complexity and scarcity of labeled data. Moreover, the causality aspects of MIP estimation are quite subjective and diverse. To this end, we aim to address the problem by annotating a large-scale ‘in-the-wild’ dataset for iden-tifying human perceptions about the ‘Most Important Person (MIP)‘ in an image. The paper provides a thorough description of our proposed Multimodal Large Language Model (MLLM) based data annotation strategy, and a thor-ough data quality analysis. Further, we perform a comprehensive benchmarking of the proposed dataset utilizing state-of-the-art MIP localization methods, indicating a significant drop in performance compared to existing datasets. The performance drop shows that the existing MIP localization algorithms must be more robust with respect to ‘in-the-wild’ situations. We believe the proposed dataset will play a vital role in building the next-generation social situation understanding methods. The dataset and associated code will be made available for research purposes.
Surbhi Madan, Shreya Ghosh 0001, Lownish Rai Sookha, M. A. Ganaie 0001, Subramanian Ramanathan, Abhinav Dhall, Tom Gedeon
WACV5
2025 Multiview Attention Fusion for Explainable Body Language Behavior Recognition
abstract
Body language behavior, including gestures and fine-grained movements not only reflects human emotions, but also serves as a versatile cue for enhancing emotional intelligence and creating responsive technologies. In this work, we explore the efficacy ofmultiview-multimodal cuesforexplainable predictionof bodily behavior. This paper proposes an attention fusion method that combines features extracted from (1) multiview videos termed “RGB”, (2) their multiview Discrete Cosine Transform representations termed “DCT” and (3) three stream skeleton features termed “Skeleton”, via a transformer-based approach. We evaluate our approach on the diverse BBSI (Balazia et al., 2022) and Drive&Act (Martin et al., 2019) datasets. Empirical results confirm that the RGB, DCT and Skeleton features enable discovery of multiple class-specific behaviors resulting in explainable predictions. Our key findings are: (a) Multimodal approaches outperform unimodal counterparts in categorizing bodily behavioral classes; (b) Efficient class predictions and plausible explanations are achieved with both unimodal and multimodal approaches; and (c) Empirical results confirm the superiority of our approach compared to state-of-the-art methods on both datasets.
Surbhi Madan, Subramanian Ramanathan, Abhinav Dhall
IEEE Trans. Affect. Comput.3
2023 A Weakly Supervised Approach to Emotion-change Prediction and Improved Mood Inference
abstract
Whilst a majority of affective computing research focuses on inferring emotions, examining mood or understanding the mood-emotion interplay has received significantly less attention. Building on prior work, we (a) deduce and incorporate emotion-change ($\Delta$) information for inferring mood, without resorting to annotated labels, and (b) attempt mood prediction for long duration video clips, in alignment with the characterisation of mood. We generate the emotion-change ($\Delta$) labels via metric learning from a pre-trained Siamese Network, and use these in addition to mood labels for mood classification. Experiments evaluating unimodal (training only using mood labels) vs muttimodat (training using mood plus $\Delta$ labels) models show that mood prediction benefits from the incorporation of emotion-change information, emphasising the importance of modelling the moodemotion interplay for effective mood inference.
Soujanya Narayana, Ibrahim Radwan, Ravikiran Parameshwara, Iman Abbasnejad, Akshay Asthana, Subramanian Ramanathan, Roland Göcke
ACII6
2023 Examining Subject-Dependent and Subject-Independent Human Affect Inference from Limited Video Data
abstract
Continuous human affect estimation from video data entails modelling the dynamic emotional state from a sequence of facial images. Though multiple affective video databases exist, they are limited in terms of data and dy-namic annotations, as assigning continuous affective labels to video data is subjective, onerous and tedious. While studies have established the existence of signature facial expressions corresponding to the basic categorical emotions, individual differences in emoting facial expressions nevertheless exist; factoring out these idiosyncrasies is critical for effective emotion inference. This work explores continuous human affect recognition using AFEW-VA, an ‘in-the-wild’ video dataset with limited data, employing subject-independent (SI) and subject-dependent (SD) settings. The SI setting involves the use of training and test sets with mutually exclusive subjects, while training and test samples corresponding to the same subject can occur in the SD setting. A novel, dynamically-weighted loss function is employed with a Convolutional Neural Network (CNN)-Long Short- Term Memory (LSTM) architecture to optimise dynamic affect prediction. Superior prediction is achieved in the SD setting, as compared to the SI counterpart.
Ravikiran Parameshwara, Ibrahim Radwan, Subramanian Ramanathan, Roland Göcke
FG3
2023 MAGIC-TBR: Multiview Attention Fusion for Transformer-based Bodily Behavior Recognition in Group Settings
abstract
Bodily behavioral language is an important social cue, and its automated analysis helps in enhancing the understanding of artificial intelligence systems. Furthermore, behavioral language cues are essential for active engagement in social agent-based user interactions. Despite the progress made in computer vision for tasks like head and body pose estimation, there is still a need to explore the detection of finer behaviors such as gesturing, grooming, or fumbling. This paper proposes a multiview attention fusion method named MAGIC-TBR that combines features extracted from videos and their corresponding Discrete Cosine Transform coefficients via a transformer-based approach. The experiments are conducted on the BBSI dataset and the results demonstrate the effectiveness of the proposed feature fusion with multiview attention. The code is available at: https://github.com/surbhimadan92/MAGIC-TBR
Surbhi Madan, Gulshan Sharma, Subramanian Ramanathan, Abhinav Dhall
ACM Multimedia4
2023 Efficient Labelling of Affective Video Datasets via Few-Shot & Multi-Task Contrastive Learning
abstract
Whilst deep learning techniques have achieved excellent emotion prediction, they still require large amounts of labelled training data, which are (a) onerous and tedious to compile, and (b) prone to errors and biases. We propose Multi-Task Contrastive Learning for Affect Representation (MT-CLAR) for few-shot affect inference. MT-CLAR combines multi-task learning with a Siamese network trained via contrastive learning to infer from a pair of expressive facial images (a) the (dis)similarity between the facial expressions, and (b) the difference in valence and arousal levels of the two faces. We further extend the image-based MT-CLAR framework for automated video labelling where, given one or a few labelled video frames (termed support-set), MT-CLAR labels the remainder of the video for valence and arousal. Experiments are performed on the AFEW-VA dataset with multiple support-set configurations; moreover, supervised learning on representations learnt via MT-CLAR are used for valence, arousal and categorical emotion prediction on the AffectNet and AFEW-VA datasets. The results show that valence and arousal predictions via MT-CLAR are very comparable to the state-of-the-art (SOTA), and we significantly outperform SOTA with a support-set ≈6% the size of the video dataset.
Ravikiran Parameshwara, Ibrahim Radwan, Akshay Asthana, Iman Abbasnejad, Subramanian Ramanathan, Roland Göcke
ACM Multimedia5
2022 Music Identification Using Brain Responses to Initial Snippets
abstract
Naturalistic music typically contains repetitive musical patterns that are present throughout the song. These patterns form a signature, enabling effortless song recognition. We investigate whether neural responses corresponding to these repetitive patterns also serve as a signature, enabling recognition of later song segments on learning initial segments. We examine EEG encoding of naturalistic musical patterns employing the NMED-T and MUSIN-G datasets. Experiments reveal that (a) training machine learning classifiers on the initial 20s song segment enables accurate prediction of the song from the remaining segments; (b) β and γ band power spectra achieve optimal song classification, and (c) listener-specific EEG responses are observed for the same stimulus, characterizing individual differences in music perception.
Pankaj Pandey, Gulshan Sharma, Krishna P. Miyapuram, Subramanian Ramanathan, Derek Lomas
ICASSP4
2022 Neural Encoding of Songs is Modulated by Their Enjoyment
abstract
We examine user and song identification from neural (EEG) signals. Owing to perceptual subjectivity in human-media interaction, music identification from brain signals is a challenging task. We demonstrate that subjective differences in music perception aid user identification, but hinder song identification. In an attempt to address intrinsic complexities in music identification, we provide empirical evidence on the role of enjoyment in song recognition. Our findings reveal that considering song enjoyment as an additional factor can improve EEG-based song recognition.
Gulshan Sharma, Pankaj Pandey, Subramanian Ramanathan, Krishna P. Miyapuram, Abhinav Dhall
ICMI3
2022 A Transformer Based Approach for Activity Detection
abstract
Non-invasive physiological sensors allow for the collection of user-specific data in realistic environments. In this paper, using physiological data, we investigate the effectiveness of Convolutional Neural Network (CNN) based feature embeddings and Transformer architecture for the human activity recognition task. 1D-CNN representation is used for the heart rate, and 2D-CNN is used for short-term Fourier transformation of the accelerometer data. Post fusion, the feature is input into a transformer. The experiments are performed on the harAGE dataset. The findings indicate the discriminative ability of the feature-fusion on transformer-based architecture, and the method outperforms the harAGE baseline by an absolute 3.7%.
Gulshan Sharma, Abhinav Dhall, Subramanian Ramanathan
ACM Multimedia3
2022 Recognition of Advertisement Emotions With Application to Computational Advertising
abstract
Advertisements (ads) often contain strong emotions to capture audience attention and convey an effective message. Still, little work has focused on affect recognition (AR) from ads employing audiovisual or user cues. This work (1) compiles an affective video ad dataset which evokes coherent emotions across users; (2) explores the efficacy of content-centric convolutional neural network (CNN) features for ad AR vis-ã-vis handcrafted audio-visual descriptors; (3) examines user-centric ad AR from Electroencephalogram (EEG) signals, and (4) demonstrates how better affect predictions facilitate effective computational advertising via a study involving 18 users. Experiments reveal that (a) CNN features outperform handcrafted audiovisual descriptors for content-centric AR; (b) EEG features encode ad-induced emotions better than content-based features; (c) Multi-task learning achieves optimal ad AR among a slew of classifiers and (d) Pursuant to (b), EEG features enable optimized ad insertion onto streamed video compared to content-based or manual insertion, maximizing ad recall and viewing experience.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan
IEEE Trans. Affect. Comput.6
2021 Head Matters: Explainable Human-centered Trait Prediction from Head Motion Dynamics
abstract
We demonstrate the utility of elementary head-motion units termed kinemes for behavioral analytics to predict personality and interview traits. Transforming head-motion patterns into a sequence of kinemes facilitates discovery of latent temporal signatures characterizing the targeted traits, thereby enabling both efficient and explainable trait prediction. Utilizing Kinemes and Facial Action Coding System (FACS) features to predict (a) OCEAN personality traits on the First Impressions Candidate Screening videos, and (b) Interview traits on the MIT dataset, we note that: (1) A Long-Short Term Memory (LSTM) network trained with kineme sequences performs better than or similar to a Convolutional Neural Network (CNN) trained with facial images; (2) Accurate predictions and explanations are achieved on combining FACS action units (AUs) with kinemes, and (3) Prediction performance is affected by the time-length over which head and facial movements are observed.
Surbhi Madan, Monika Gahalawat, Tanaya Guha, Subramanian Ramanathan
ICMI4
2021 ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction
abstract
We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a saliency map via trilinear interpolation and 3D convolutions, combining features from multiple hierarchies. The overall architecture of ViNet is conceptually simple; it is causal and runs in real-time (60 fps). ViNet does not use audio as input and still outperforms the state-of-the-art audio-visual saliency prediction models on nine different datasets (three visual-only and six audio-visual datasets). ViNet also surpasses human performance on the CC, SIM and AUC metrics for the AVE dataset, and to our knowledge, it is the first model to do so. We also explore a variation of ViNet architecture by augmenting audio features into the decoder. To our surprise, upon sufficient training, the network becomes agnostic to the input audio and provides the same output irrespective of the input. Interestingly, we also observe similar behaviour in the previous state-of-the-art models [1] for audio-visual saliency prediction. Our findings contrast with previous works on deep learning-based audio-visual saliency prediction, suggesting a clear avenue for future explorations incorporating audio in a more effective manner. The code and pre-trained models are available at https://github.com/samyak0210/ViNet.
Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti, Shyamgopal Karthik, Subramanian Ramanathan, Vineet Gandhi
IROS5
2020 GAZED- Gaze-guided Cinematic Editing of Wide-Angle Monocular Video Recordings
abstract
We present GAZED– eye GAZe-guided EDiting for videos captured by a solitary, static, wide-angle and high-resolution camera. Eye-gaze has been effectively employed in computational applications as a cue to capture interesting scene content; we employ gaze as a proxy to select shots for inclusion in the edited video. Given the original video, scene content and user eye-gaze tracks are combined to generate an edited video comprising cinematically valid actor shots and shot transitions to generate an aesthetic and vivid representation of the original narrative. We model cinematic video editing as an energy minimization problem over shot selection, whose constraints capture cinematographic editing conventions. Gazed scene locations primarily determine the shots constituting the edited video. Effectiveness of GAZED against multiple competing methods is demonstrated via a psychophysical study involving 12 users and twelve performance videos.
K. L. Bhanu Moorthy, Moneish Kumar, Subramanian Ramanathan, Vineet Gandhi
CHI3
2020 The eyes know it: FakeET- An Eye-tracking Database to Understand Deepfake Perception
abstract
We present FakeET -- an eye-tracking database to understand human visual perception of deepfake videos. Given that the principal purpose of deepfakes is to deceive human observers, FakeET is designed to understand and evaluate the ability of viewers to detect synthetic video artifacts. FakeET contains viewing patterns compiled from 40 users via the Tobii desktop eye-tracker for 811 videos from the Google Deepfake dataset, with a minimum of two viewings per video. Additionally, EEG responses acquired via the Emotiv sensor are also available. The compiled data confirms (a) distinct eye movement characteristics for real vs fake videos; (b) utility of the eye-track saliency maps for spatial forgery localization and detection, and (c) Error Related Negativity (ERN) triggers in the EEG responses, and the ability of the raw EEG signal to distinguish between real and fake videos.
Komal Chugh, Abhinav Dhall, Subramanian Ramanathan
ICMI4
2020 Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization
abstract
We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to dis-harmony between the two modalities, e.g., loss of lip-sync, unnatural facial and lip movements, etc. MDS is computed as the mean aggregate of dissimilarity scores between audio and visual segments in a video. Discriminative features are learnt for the audio and visual channels in a chunk-wise manner, employing the cross-entropy loss for individual modalities, and a contrastive loss that models inter-modality similarity. Extensive experiments on the DFDC and DeepFake-TIMIT Datasets show that our approach outperforms the state-of-the-art by up to 7%. We also demonstrate temporal forgery localization, and show how our technique identifies the manipulated video segments.
Komal Chugh, Abhinav Dhall, Subramanian Ramanathan
ACM Multimedia4
2018 EEG-based Evaluation of Cognitive Workload Induced by Acoustic Parameters for Data Sonification
abstract
Data Visualization has been receiving growing attention recently, with ubiquitous smart devices designed to render information in a variety of ways. However, while evaluations of visual tools for their interpretability and intuitiveness have been commonplace, not much research has been devoted to other forms of data rendering, \eg, sonification. This work is the first to automatically estimate the cognitive load induced by different acoustic parameters considered for sonification in prior studies~\citeferguson2017evaluation,ferguson2018investigating. We examine cognitive load via (a) perceptual data-sound mapping accuracies of users for the different acoustic parameters, (b) cognitive workload impressions explicitly reported by users, and (c) their implicit EEG responses compiled during the mapping task. Our main findings are that (i) low cognitive load-inducing (ıe, more intuitive) acoustic parameters correspond to higher mapping accuracies, (ii) EEG spectral power analysis reveals higher α band power for low cognitive load parameters, implying a congruent relationship between explicit and implicit user responses, and (iii) Cognitive load classification with EEG features achieves a peak F1-score of 0.64, confirming that reliable workload estimation is achievable with user EEG data compiled using wearable sensors.
Maneesh Bilalpur, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan
ICMI4
2018 Looking Beyond a Clever Narrative: Visual Context and Attention are Primary Drivers of Affect in Video Advertisements
abstract
Emotion evoked by an advertisement plays a key role in influencing brand recall and eventual consumer choices. Automatic ad affect recognition has several useful applications. However, the use of content-based feature representations does not give insights into how affect is modulated by aspects such as the ad scene setting, salient object attributes and their interactions. Neither do such approaches inform us on how humans prioritize visual information for ad understanding. Our work addresses these lacunae by decomposing video content into detected objects, coarse scene structure, object statistics and actively attended objects identified via eye-gaze. We measure the importance of each of these information channels by systematically incorporating related information into ad affect prediction models. Contrary to the popular notion that ad affect hinges on the narrative and the clever use of linguistic and social cues, we find that actively attended objects and the coarse scene structure better encode affective information as compared to individual scene objects or conspicuous background elements.
Abhinav Shukla, Harish Katti, Mohan Kankanhalli, Subramanian Ramanathan
ICMI4
2018 AVEID: Automatic Video System for Measuring Engagement In Dementia
abstract
Engagement in dementia is typically measured using behavior observational scales (BOS) that are tedious and involve intensive manual labor to annotate, and are therefore not easily scalable. We propose AVEID, a low cost and easy-to-use video-based engagement measurement tool to determine the engagement level of a person with dementia (PwD) during digital interaction. We show that the objective behavioral measures computed via AVEID correlate well with subjective expert impressions for the popular MPES and OME BOS, confirming its viability and effectiveness. Moreover, AVEID measures can be obtained for a variety of engagement designs, thereby facilitating large-scale studies with PwD populations.
Viral Parekh, Pin Sym Foong, Shengdong Zhao 0001, Subramanian Ramanathan
IUI4
2018 Watch to Edit: Video Retargeting using Gaze
abstract
Abstract We present a novel approach to optimally retarget videos for varied displays with differing aspect ratios by preserving salient scene content discovered via eye tracking. Our algorithm performs editing with cut, pan and zoom operations by optimizing the path of a cropping window within the original video while seeking to (i) preserve salient regions, and (ii) adhere to the principles of cinematography. Our approach is (a) content agnostic as the same methodology is employed to re‐edit a wide‐angle video recording or a close‐up movie sequence captured with a static or moving camera, and (b) independent of video length and can in principle re‐edit an entire movie in one shot. Our algorithm consists of two steps. The first step employs gaze transition cues to detect time stamps where new cuts are to be introduced in the original video via dynamic programming. A subsequent step optimizes the cropping window path (to create pan and zoom effects), while accounting for the original and new cuts. The cropping window path is designed to include maximum gaze information, and is composed of piecewise constant, linear and parabolic segments. It is obtained via L(1)regularized convex optimization which ensures a smooth viewing experience. We test our approach on a wide variety of videos and demonstrate significant improvement over the state‐of‐the‐art, both in terms of computational complexity and qualitative aspects. A study performed with 16 users confirms that our approach results in a superior viewing experience as compared to gaze driven re‐editing [ JSSH15 ] and letterboxing methods, especially for wide‐angle static camera recordings.
Kranthi Kumar Rachavarapu, Moneish Kumar, Vineet Gandhi, Subramanian Ramanathan
Comput. Graph. Forum4
2018 Joint Estimation of Human Pose and Conversational Groups from Social Scenes
Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz, Elisa Ricci 0001
Int. J. Comput. Vis.2
2018 ASCERTAIN: Emotion and Personality Recognition Using Commercial Sensors
abstract
We present ASCERTAIN-a multimodal databaASe for impliCit pERsonaliTy and Affect recognitIoN using commercial physiological sensors. To our knowledge, ASCERTAIN is the first database to connect personality traits and emotional states via physiological responses. ASCERTAIN contains big-five personality scales and emotional self-ratings of 58 users along with their Electroencephalogram (EEG), Electrocardiogram (ECG), Galvanic Skin Response (GSR) and facial activity data, recorded using off-the-shelf sensors while viewing affective movie clips. We first examine relationships between users' affective ratings and personality scales in the context of prior observations, and then study linear and non-linear physiological correlates of emotion and personality. Our analysis suggests that the emotion-personality relationship is better captured by non-linear rather than linear statistics. We finally attempt binary emotion and personality trait recognition using physiological features. Experimental results cumulatively confirm that personality differences are better revealed while comparing user responses to emotionally homogeneous videos, and above-chance recognition is achieved for both affective and personality dimensions.
Subramanian Ramanathan, Julia Wache, Mojtaba Khomami Abadi, Radu L. Vieriu, Stefan Winkler 0001, Nicu Sebe
IEEE Trans. Affect. Comput.1
2017 Discovering gender differences in facial emotion recognition via implicit behavioral cues
abstract
We examine the utility of implicit behavioral cues in the form of EEG brain signals and eye movements for gender recognition (GR) and emotion recognition (ER). Specifically, the examined cues are acquired via low-cost, off-the-shelf sensors. We asked 28 viewers (14 female) to recognize emotions from unoccluded (no mask) as well as partially occluded (eye and mouth masked) emotive faces. Obtained experimental results reveal that (a) reliable GR and ER is achievable with EEG and eye features, (b) differential cognitive processing especially for negative emotions is observed for males and females and (c) some of these cognitive differences manifest under partial face occlusion, as typified by the eye and mouth mask conditions.
Maneesh Bilalpur, Seyed Mostafa Kia, Tat-Seng Chua, Subramanian Ramanathan
ACII4
2017 Gender and emotion recognition with implicit user signals
abstract
We examine the utility of implicit user behavioral signals captured using low-cost, off-the-shelf devices for anonymous gender and emotion recognition. A user study designed to examine male and female sensitivity to facial emotions confirms that females recognize (especially negative) emotions quicker and more accurately than men, mirroring prior findings. Implicit viewer responses in the form of EEG brain signals and eye movements are then examined for existence of (a) emotion and gender-specific patterns from event-related potentials (ERPs) and fixation distributions and (b) emotion and gender discriminability. Experiments reveal that (i) Gender and emotion-specific differences are observable from ERPs, (ii) multiple similarities exist between explicit responses gathered from users and their implicit behavioral signals, and (iii) Significantly above-chance (≈70%) gender recognition is achievable on comparing emotion-specific EEG responses– gender differences are encoded best for anger and disgust. Also, fairly modest valence (positive vs negative emotion) recognition is achieved with EEG and eye-based features.
Maneesh Bilalpur, Seyed Mostafa Kia, Manisha Chawla, Tat-Seng Chua, Subramanian Ramanathan
ICMI5
2017 Evaluating content-centric vs. user-centric ad affect recognition
abstract
Despite the fact that advertisements (ads) often include strongly emotional content, very little work has been devoted to affect recognition (AR) from ads. This work explicitly compares content-centric and user-centric ad AR methodologies, and evaluates the impact of enhanced AR on computational advertising via a user study. Specifically, we (1) compile an affective ad dataset capable of evoking coherent emotions across users; (2) explore the efficacy of content-centric convolutional neural network (CNN) features for encoding emotions, and show that CNN features outperform low-level emotion descriptors; (3) examine user-centered ad AR by analyzing Electroencephalogram (EEG) responses acquired from eleven viewers, and find that EEG signals encode emotional information better than content descriptors; (4) investigate the relationship between objective AR and subjective viewer experience while watching an ad-embedded online video stream based on a study involving 12 users. To our knowledge, this is the first work to (a) expressly compare user vs content-centered AR for ads, and (b) study the relationship between modeling of ad emotions and its impact on a real-life advertising application.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Subramanian Ramanathan
ICMI6
2017 Affect Recognition in Ads with Application to Computational Advertising
abstract
Advertisements (ads) often include strongly emotional content to leave a lasting impression on the viewer. This work (i) compiles an affective ad dataset capable of evoking coherent emotions across users, as determined from the affective opinions of five experts and 14 annotators; (ii) explores the efficacy of convolutional neural network (CNN) features for encoding emotions, and observes that CNN features outperform low-level audio-visual emotion descriptors[9] upon extensive experimentation; and (iii) demonstrates how enhanced affect prediction facilitates computational advertising, and leads to better viewing experience while watching an online video stream embedded with ads based on a study involving 17 users. We model ad emotions based on subjective human opinions as well as objective multimodal features, and show how effectively modeling ad emotions can positively impact a real-life application.
Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Karthik Yadati, Mohan Kankanhalli, Subramanian Ramanathan
ACM Multimedia6
2017 Active Online Anomaly Detection Using Dirichlet Process Mixture Model and Gaussian Process Classification
abstract
We present a novel anomaly detection (AD) system for streaming videos. Different from prior methods that rely on unsupervised learning of clip representations, that are usually coarse in nature, and batch-mode learning, we propose the combination of two non-parametric models for our task: (i) Dirichlet process mixture models (DPMM) based modeling of object motion and directions in each cell, and (ii) Gaussian process based active learning paradigm involving labeling by a domain expert. Whereas conventional clip representation methods adopt quantizing only motion directions leading to a lossy, coarse representation that are inadequate, our clip representation approach results in fine grained clusters at each cell that model the scene activities (both direction and speed) more effectively. For active anomaly detection, we adapt a Gaussian Process framework to process incoming samples (video snippets) sequentially, seek labels for confusing or informative samples and and update the AD model online. Furthermore, the proposed video representation along with a novel query criterion to select informative samples for labeling that incorporates both exploration and exploitation criteria is proposed, and is found to outperform competing criteria on two challenging traffic scene datasets.
Jagannadan Varadarajan, Subramanian Ramanathan, Narendra Ahuja, Pierre Moulin, Jean-Marc Odobez
WACV2
2017 A Probabilistic Approach to People-Centric Photo Selection and Sequencing
abstract
We present a crowdsourcing (CS) study to examine how specific attributes probabilistically affect the selection and sequencing of images from personal photo collections. Thirteen image attributes are explored, including seven people-centric properties. We first propose a novel dataset shaping technique based on mixed integer linear programming (MILP) to identify a subset of photos in which the attributes of interest are uniformly distributed and minimally correlated. Shaping enables the synthesis of compact, balanced, and representative datasets for CS, and facilitates effective learning of the selection likelihood of an image as well as its relative position in a sequence, given its attributes. We further present an ILP-based slideshow creation framework to select and arrange (a subset of) appealing images from a personal photo library. Quantitative and qualitative evaluations confirm that our method outperforms regression-based and greedy approaches for photo selection and sequencing, generating slideshows similar in quality to those created by humans.
Vassilios Vonikakis, Subramanian Ramanathan, Jonas Toft Arnfred, Stefan Winkler 0001
IEEE Trans. Multim.2
2016 State-Action Based Link Layer Design for IEEE 802.11b Compliant MATLAB-Based SDR
abstract
Software defined radio (SDR) allows unprecedented levels of flexibility by transitioning the radio communication system from a rigid hardware platform to a more user-controlled software paradigm. However, it can still be time consuming to design and implement such SDRs as they typically require thorough knowledge of the operating environment and a careful tuning of the program. In this work, we describe a systems contribution and outline strategies on how to create a state-action based design in implementing the CSMA/CA/ACK MAC layer in MATLAB®that runs on the USRP®platform, a commonly used SDR. Our design allows optimal selection of the parameters so that all operations remain functionally compliant with the IEEE 802.11b standard (1Mbps specification). The code base of the system is enabled through the Communications System ToolboxTMand incorporates channel sensing and exponential random back-off for contention resolution. The current work provides a testbed to experiment with and enables creation of new MAC protocols starting from the fundamental IEEE 802.11b compliant standard. Our system design approach guarantees the consistent performance of the bi-directional link and we include the experimental results for the three node system to demonstrate the robustness of the MAC layer in mitigating packet collisions and enforcing fairness among nodes.
Subramanian Ramanathan, Eric Doyle, Benjamin Drozdenko, Miriam Leeser, Kaushik R. Chowdhury
DCOSS1
2016 Shaping datasets: Optimal data selection for specific target distributions across dimensions
abstract
This paper presents a method for dataset manipulation based on Mixed Integer Linear Programming (MILP). The proposed optimization can narrow down a dataset to a particular size, while enforcing specific distributions across different dimensions. It essentially leverages the redundancies of an initial dataset in order to generate more compact versions of it, with a specific target distribution across each dimension. If the desired target distribution is uniform, then the effect is balancing: all values across all different dimensions are equally represented. Other types of target distributions can also be specified, depending on the nature of the problem. The proposed approach may be used in machine learning, for shaping training and testing datasets, or in crowdsourcing, for preparing datasets of a manageable size.
Vassilios Vonikakis, Subramanian Ramanathan, Stefan Winkler 0001
ICIP2
2016 COVERAGE - A novel database for copy-move forgery detection
abstract
We present COVERAGE - a novel database containing copy-move forged images and their originals with similar but genuine objects. COVERAGE is designed to highlight and address tamper detection ambiguity of popular methods, caused by self-similarity within natural images. In COVERAGE, forged-original pairs are annotated with (i) the duplicated and forged region masks, and (ii) the tampering factor/similarity metric. For benchmarking, forgery quality is evaluated using (i) computer vision-based methods, and (ii) human detection performance. We also propose a novel sparsity-based metric for efficiently estimating forgery quality. Experimental results show that (a) popular forgery detection methods perform poorly over COVERAGE, and (b) the proposed sparsity based metric best correlates with human detection performance. We release the COVERAGE database to the research community.
Bihan Wen, Subramanian Ramanathan, Tian-Tsong Ng, Xuanjing Shen, Stefan Winkler 0001
ICIP3
2016 SALSA: A Novel Dataset for Multimodal Group Behavior Analysis
abstract
Studying free-standing conversational groups (FCGs) in unstructured social settings (e.g., cocktail party ) is gratifying due to the wealth of information available at the group (mining social networks) and individual (recognizing native behavioral and personality traits) levels. However, analyzing social scenes involving FCGs is also highly challenging due to the difficulty in extracting behavioral cues such as target locations, their speaking activity and head/body pose due to crowdedness and presence of extreme occlusions. To this end, we propose SALSA, a novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis, and make two main contributions to research on automated social interaction analysis: (1) SALSA records social interactions among 18 participants in a natural, indoor environment for over 60 minutes, under the poster presentation and cocktail party contexts presenting difficulties in the form of low-resolution images, lighting variations, numerous occlusions, reverberations and interfering sound sources; (2) To alleviate these problems we facilitate multimodal analysis by recording the social interplay using four static surveillance cameras and sociometric badges worn by each participant, comprising the microphone, accelerometer, bluetooth and infrared sensors. In addition to raw data, we also provide annotations concerning individuals' personality as well as their position, head, body orientation and F-formation information over the entire event duration. Through extensive experiments with state-of-the-art approaches, we show (a) the limitations of current methods and (b) how the recorded multiple cues synergetically aid automatic analysis of social interactions. SALSA is available at http://tev.fbk.eu/salsa.
Xavier Alameda-Pineda, Jacopo Staiano, Subramanian Ramanathan, Ligia Maria Batrinca, Elisa Ricci 0001, Bruno Lepri, Oswald Lanz, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 A Multi-Task Learning Framework for Head Pose Estimation under Target Motion
abstract
Recently, head pose estimation (HPE) from low-resolution surveillance data has gained in importance. However, monocular and multi-view HPE approaches still work poorly under target motion, as facial appearance distorts owing to camera perspective and scale changes when a person moves around. To this end, we propose FEGA-MTL, a novel framework based on Multi-Task Learning (MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. Upon partitioning the monitored scene into a dense uniform spatial grid, FEGA-MTL simultaneously clusters grid partitions into regions with similar facial appearance, while learning region-specific head pose classifiers. In the learning phase, guided by two graphs which a-priori model the similarity among (1) grid partitions based on camera geometry and (2) head pose classes, FEGA-MTL derives the optimal scene partitioning and associated pose classifiers. Upon determining the target's position using a person tracker at test time, the corresponding region-specific classifier is invoked for HPE. The FEGA-MTL framework naturally extends to a weakly supervised setting where the target's walking direction is employed as a proxy in lieu of head orientation. Experiments confirm that FEGA-MTL significantly outperforms competing single-task and multi-task learning methods in multi-view settings.
Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Gaowen Liu, Oswald Lanz, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Active domain adaptation with noisy labels for multimedia analysis
Gaowen Liu, Yan Yan 0002, Subramanian Ramanathan, Jingkuan Song, Guoyu Lu 0001, Nicu Sebe
World Wide Web3
2015 Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations from Surveillance Videos
abstract
We present a novel approach for jointly estimating targets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by exploiting the limited range of orientations that the two can jointly take, or (ii) determined F-formations based on the mutual head (but not body) orientations of interactors, we present a unified framework to jointly infer both (i) and (ii). Apart from exploiting spatial and orientation relationships, we also integrate cues pertaining to temporal consistency and occlusions, which are beneficial while handling low-resolution data under surveillance settings. Efficacy of the joint inference framework reflects via increased head, body pose and F-formation estimation accuracy over the state-of-the-art, as confirmed by extensive experiments on two social datasets.
Elisa Ricci 0001, Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz
ICCV3
2015 On the utility of canonical correlation analysis for domain adaptation in multi-view headpose estimation
abstract
The utility of canonical correlation analysis (CCA) for domain adaptation (DA) in the context of multi-view head pose estimation is examined in this work. We consider the three problems studied in [1], where different DA approaches are explored to transfer head pose-related knowledge from an extensively labeled source dataset to a sparsely labeled target set, whose attributes are vastly different from the source. CCA is found to benefit DA for all the three problems, and the use of a covariance profile-based diagonality score (DS) also improves classification performance with respect to a nearest neighbor (NN) classifier.
Anoop Kolar Rajagopal, Subramanian Ramanathan, Vassilios Vonikakis, K. R. Ramakrishnan, Stefan Winkler 0001
ICIP2
2015 PET: An eye-tracking dataset for animal-centric Pascal object classes
abstract
We present PET- the Pascal animal classes Eye Tracking database. Our database comprises eye movement recordings compiled from forty users for the bird, cat, cow, dog, horse and sheep trainval sets from the VOC 2012 image set. Different from recent eye-tracking databases such as [1, 2], a salient aspect of PET is that it contains eye movements recorded for both the free-viewing and visual search task conditions. While some differences in terms of overall gaze behavior and scanning patterns are observed between the two conditions, a very similar number of fixations are observed on target objects for both conditions. As a utility application, we show how feature pooling around fixated locations enables enhanced (animal) object classification accuracy.
Syed Omer Gilani, Subramanian Ramanathan, Yan Yan 0002, David Melcher, Nicu Sebe, Stefan Winkler 0001
ICME2
2015 Implicit User-centric Personality Recognition Based on Physiological Responses to Emotional Videos
abstract
We present a novel framework for recognizing personality traits based on users' physiological responses to affective movie clips. Extending studies that have correlated explicit/implicit affective user responses with Extraversion and Neuroticism traits, we perform single-trial recognition of the big-five traits from Electrocardiogram (ECG), Galvanic Skin Response (GSR), Electroencephalogram (EEG) and facial emotional responses compiled from 36 users using off-the-shelf sensors. Firstly, we examine relationships among personality scales and (explicit) affective user ratings acquired in the context of prior observations. Secondly, we isolate physiological correlates of personality traits. Finally, unimodal and multimodal personality recognition results are presented. Personality differences are better revealed while analyzing responses to emotionally homogeneous (e.g., high valence, high arousal) clips, and significantly above-chance recognition is achieved for all five traits.
Julia Wache, Subramanian Ramanathan, Mojtaba Khomami Abadi, Radu L. Vieriu, Nicu Sebe, Stefan Winkler 0001
ICMI2
2015 Jointly Estimating Interactions and Head, Body Pose of Interactors from Distant Social Scenes
abstract
We present joint estimation of F-formations and head, body pose of interactors in a social scene captured by surveillance cameras. Unlike prior works that have focused on (a) discovering F-formations based on head pose and position cues, or (b) jointly learned head and body pose of individuals based on anatomic constraints, we exploit positional and pose cues characterizing interactors and interactions to jointly infer both (a) and (b). We show how the joint inference framework benefits both F-formation and head, body pose estimation accuracy via experiments on two social datasets.
Subramanian Ramanathan, Jagannadan Varadarajan, Elisa Ricci 0001, Oswald Lanz, Stefan Winkler 0001
ACM Multimedia1
2015 DECAF: MEG-Based Multimodal Database for Decoding Affective Physiological Responses
abstract
In this work, we present DECAF-a multimodal data set for decoding user physiological responses to affective multimedia content. Different from data sets such as DEAP [15] and MAHNOB-HCI [31], DECAF contains (1) brain signals acquired using the Magnetoencephalogram (MEG) sensor, which requires little physical contact with the user's scalp and consequently facilitates naturalistic affective response, and (2) explicit and implicit emotional responses of 30 participants to 40 one-minute music video segments used in [15] and 36 movie clips, thereby enabling comparisons between the EEG versus MEG modalities as well as movie versus music stimuli for affect recognition. In addition to MEG data, DECAF comprises synchronously recorded near-infra-red (NIR) facial videos, horizontal Electrooculogram (hEOG), Electrocardiogram (ECG), and trapezius-Electromyogram (tEMG) peripheral physiological responses. To demonstrate DECAF's utility, we present (i) a detailed analysis of the correlations between participants' self-assessments and their physiological responses and (ii) single-trial classification results for valence, arousal and dominance, with performance evaluation against existing data sets. DECAF also contains time-continuous emotion annotations for movie clips from seven users, which we use to demonstrate dynamic emotion prediction.
Mojtaba Khomami Abadi, Subramanian Ramanathan, Seyed Mostafa Kia, Paolo Avesani, Ioannis Patras, Nicu Sebe
IEEE Trans. Affect. Comput.2
2014 WADE: simplified GUI add-on development for third-party software
abstract
We present the WADE Integrated Development Environment (IDE), which simplifies interface and functionality modification of existing third-party software without access to source code. WADE clones the Graphical User Interface (GUI) of a host program through dynamic-link library (DLL) injection, enabling modifications to (1) the GUI in a WYSIWYG fashion and (2) software functionality. We compare WADE with an alternative state-of-the-art runtime toolkit overloading approach in a user-study, whose results demonstrate that WADE significantly simplifies the task of GUI-based add-on development.
Xiaojun Meng, Shengdong Zhao 0001, James R. Eagan, Subramanian Ramanathan
CHI6
2014 Clustered Multi-task Linear Discriminant Analysis for View Invariant Color-Depth Action Recognition
abstract
The widespread adoption of low-cost depth cameras has opened new opportunities to improve traditional action recognition systems. In this paper we focus on the specific problem of action recognition under view point changes and propose a novel approach for view-invariant action recognition operating jointly on visual data of color and depth camera channels. Our method is based on the unique combination of robust Self-Similarity Matrix (SSM) descriptors and multi-task learning. Indeed, multi-view action recognition is inherently a multi-task learning problem: images from a camera view can be modeled as visual data associated to the same task and it is reasonable to assume that the data of different tasks (camera views) are related to each other. In this work we propose a novel algorithm extending Multi-Task Linear Discriminant Analysis (MT-LDA) to enhance its flexibility by learning the dependencies between different views. Extensive experimental results on the publicly available ACT42dataset demonstrate the effectiveness of the proposed method.
Yan Yan 0002, Elisa Ricci 0001, Gaowen Liu, Subramanian Ramanathan, Nicu Sebe
ICPR4
2014 Evaluating Multi-task Learning for Multi-view Head-Pose Classification in Interactive Environments
abstract
Social attention behavior offers vital cues towards inferring one's personality traits from interactive settings such as round-table meetings and cocktail parties. Head orientation is typically employed as a proxy for determining the social attention direction when faces are captured at low-resolution. Recently, multi-task learning has been proposed to robustly compute head pose under perspective and scale-based facial appearance variations when multiple, distant and large field-of-view cameras are employed for visual analysis in smart-room applications. In this paper, we evaluate the effectiveness of an SVM-based MTL (SVM+MTL) framework with various facial descriptors (KL, HOG, LBP, etc.). The KL+HOG feature combination is found to produce the best classification performance, with SVM+MTL outperforming classical SVM irrespective of the feature used.
Yan Yan 0002, Subramanian Ramanathan, Elisa Ricci 0001, Oswald Lanz, Nicu Sebe
ICPR2
2014 Exploring Transfer Learning Approaches for Head Pose Classification from Multi-view Surveillance Images
Anoop Kolar Rajagopal, Subramanian Ramanathan, Elisa Ricci 0001, Radu L. Vieriu, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe
Int. J. Comput. Vis.2
2014 Learning from multiple annotators with varying expertise
Yan Yan 0024, Rómer Rosales, Glenn Fung, Subramanian Ramanathan, Jennifer G. Dy
Mach. Learn.4
2014 Multitask Linear Discriminant Analysis for View Invariant Action Recognition
abstract
Robust action recognition under viewpoint changes has received considerable attention recently. To this end, self-similarity matrices (SSMs) have been found to be effective view-invariant action descriptors. To enhance the performance of SSM-based methods, we propose multitask linear discriminant analysis (LDA), a novel multitask learning framework for multiview action recognition that allows for the sharing of discriminative SSM features among different views (i.e., tasks). Inspired by the mathematical connection between multivariate linear regression and LDA, we model multitask multiclass LDA as a single optimization problem by choosing an appropriate class indicator matrix. In particular, we propose two variants of graph-guided multitask LDA: 1) where the graph weights specifying view dependencies are fixed a priori and 2) where graph weights are flexibly learnt from the training data. We evaluate the proposed methods extensively on multiview RGB and RGBD video data sets, and experimental results confirm that the proposed approaches compare favorably with the state-of-the-art.
Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Gaowen Liu, Nicu Sebe
IEEE Trans. Image Process.3
2013 User-centric Affective Video Tagging from MEG and Peripheral Physiological Responses
abstract
This paper presents a new multimodal database and the associated results for characterization of affect (valence, arousal and dominance) using the Magneto encephalogram (MEG) brain signals and peripheral physiological signals (horizontal EOG, ECG, trapezius EMG). We attempt single-trial classification of affect in movie and music video clips employing emotional responses extracted from eighteen participants. The main findings of this study are that: (i) the MEG signal effectively encodes affective viewer responses, (ii) clip arousal is better predicted by MEG, while peripheral physiological signals are more effective for predicting valence and (iii) prediction performance is better for movie clips as compared to music video clips.
Mojtaba Khomami Abadi, Seyed Mostafa Kia, Subramanian Ramanathan, Paolo Avesani, Nicu Sebe
ACII3
2013 No Matter Where You Are: Flexible Graph-Guided Multi-task Learning for Multi-view Head Pose Classification under Target Motion
abstract
We propose a novel Multi-Task Learning framework (FEGA-MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. As the target (person) moves, distortions in facial appearance owing to camera perspective and scale severely impede performance of traditional head pose classification methods. FEGA-MTL operates on a dense uniform spatial grid and learns appearance relationships across partitions as well as partition-specific appearance variations for a given head pose to build region-specific classifiers. Guided by two graphs which a-priori model appearance similarity among (i) grid partitions based on camera geometry and (ii) head pose classes, the learner efficiently clusters appearance wise related grid partitions to derive the optimal partitioning. For pose classification, upon determining the target's position using a person tracker, the appropriate region specific classifier is invoked. Experiments confirm that FEGA-MTL achieves state-of-the-art classification with few training data.
Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Oswald Lanz, Nicu Sebe
ICCV3
2013 Impact of image appeal on visual attention during photo triaging
abstract
Image appeal is determined by factors such as exposure, white balance, motion blur, scene perspective, and semantics. All these factors influence the selection of the best image(s) in a typical photo triaging task. This paper presents the results of an exploratory study on how image appeal affected selection behavior and visual attention patterns of 11 users, who were assigned the task of selecting the best photo from each of 40 groups. Images with low appeal were rejected, while highly appealing images were selected by a majority. Images with higher appeal attracted more visual attention, and users spent more time exploring them. A comparison of user eye fixation maps with three state-of-the-art saliency models revealed that these differences are not captured by the models.
Syed Omer Gilani, Subramanian Ramanathan, Huang Hua, Stefan Winkler 0001, Shih-Cheng Yen
ICIP2
2013 On the relationship between head pose, social attention and personality prediction for unstructured and dynamic group interactions
abstract
Correlates between social attention and personality traits have been widely acknowledged in social psychology studies. Head pose has commonly been employed as a proxy for determining the social attention direction in small group interactions. However, the impact of head pose estimation errors on personality estimates has not been studied to our knowledge.
Subramanian Ramanathan, Yan Yan 0002, Jacopo Staiano, Oswald Lanz, Nicu Sebe
ICMI1
2012 An Adaptation Framework for Head-Pose Classification in Dynamic Multi-view Scenarios
Anoop Kolar Rajagopal, Subramanian Ramanathan, Radu L. Vieriu, Elisa Ricci 0001, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe
ACCV (2)2
2012 Active transfer learning for multi-view head-pose classification
Yan Yan 0002, Subramanian Ramanathan, Oswald Lanz, Nicu Sebe
ICPR2
2012 Connecting Meeting Behavior with Extraversion - A Systematic Study
abstract
This work investigates the suitability of medium-grained meeting behaviors, namely, speaking time and social attention, for automatic classification of the Extraversion personality trait. Experimental results confirm that these behaviors are indeed effective for the automatic detection of Extraversion. The main findings of our study are that: 1) Speaking time and (some forms of) social gaze are effective indicators of Extraversion, 2) classification accuracy is affected by the amount of time for which meeting behavior is observed, 3) independently considering only the attention received by the target from peers is insufficient, and 4) distribution of social attention of peers plays a crucial role.
Bruno Lepri, Subramanian Ramanathan, Kyriaki Kalimeri, Jacopo Staiano, Fabio Pianesi, Nicu Sebe
IEEE Trans. Affect. Comput.2
2011 Can computers learn from humans to see better?: inferring scene semantics from viewers' eye movements
abstract
This paper describes an attempt to bridge the semantic gap between computer vision and scene understanding employing eye movements. Even as computer vision algorithms can efficiently detect scene objects, discovering semantic relationships between these objects is as essential for scene understanding. Humans understand complex scenes by rapidly moving their eyes (saccades) to selectively focus on salient entities (fixations). For 110 social scenes, we compared verbal descriptions provided by observers against eye movements recorded during a free-viewing task. Data analysis confirms (i) a strong correlation between task-explicit linguistic descriptions and task-implicit eye movements, both of which are influenced by underlying scene semantics and (ii) the ability of eye movements in the form of fixations and saccades to indicate salient entities and entity relationships mentioned in scene descriptions.
Subramanian Ramanathan, Victoria Yanulevskaya, Nicu Sebe
ACM Multimedia1
2011 Automatic modeling of personality states in small group interactions
abstract
In this paper, we target the automatic recognition of personality states in a meeting scenario employing visual and acoustic features. The social psychology literature has coined the name personality state to refer to a specific behavioral episode wherein a person behaves as more or less introvert/extrovert, neurotic or open to experience, etc. Personality traits can then be reconstructed as density distributions over personality states. Different machine learning approaches were used to test the effectiveness of the selected features in modeling the dynamics of personality states.
Jacopo Staiano, Bruno Lepri, Subramanian Ramanathan, Nicu Sebe, Fabio Pianesi
ACM Multimedia3
2010 An Eye Fixation Database for Saliency Detection in Images
Subramanian Ramanathan, Harish Katti, Nicu Sebe, Mohan Kankanhalli, Tat-Seng Chua
ECCV (4)1
2010 Making computers look the way we look: exploiting visual attention for image understanding
abstract
Human Visual attention (HVA) is an important strategy to focus on specific information while observing and understanding visual stimuli. HVA involves making a series of fixations on select locations while performing tasks such as object recognition, scene understanding, etc. We present one of the first works that combines fixation information with automated concept detectors to (i) infer abstract image semantics, and (ii) enhance performance of object detectors.
Harish Katti, Subramanian Ramanathan, Mohan Kankanhalli, Nicu Sebe, Tat-Seng Chua, K. R. Ramakrishnan
ACM Multimedia2
2010 Putting the pieces together: multimodal analysis of social attention in meetings
abstract
This paper presents a multimodal framework employing eye-gaze, head-pose and speech cues to explain observed social attention patterns in meeting scenes. We first investigate a few hypotheses concerning social attention and characterize meetings and individuals based on ground-truth data. This is followed by replication of ground-truth results through automated estimation of eye-gaze, head-pose and speech activity for each participant. Experimental results show that combining eye-gaze and head-pose estimates decreases error in social attention estimation by over 26%.
Subramanian Ramanathan, Jacopo Staiano, Kyriaki Kalimeri, Nicu Sebe, Fabio Pianesi
ACM Multimedia1
2009 A robust framework for aligning lecture slides with video
abstract
We propose a robust approach for aligning lecture slides with lecture videos using a combination of Hough transform, optical flow and Gabor analysis. A Markov Decision Process model is used to incorporate prior knowledge for enhanced recognition. We demonstrate synchronization of slides with videos containing de-focused slide content, speaker occlusion as well as camera pan, tilt and zoom sequences. Experimental results confirm the effectiveness of our approach for multimedia indexing applications.
Xiangyu Wang 0002, Subramanian Ramanathan, Mohan Kankanhalli
ICIP2
2009 Automated localization of affective objects and actions in images via caption text-cum-eye gaze analysis
abstract
We propose a novel framework to localize and label affective objects and actions in images through a combination of text, visual and gaze-based analysis. Human gaze provides useful cues to infer locations and interactions of affective objects. While concepts (labels) associated with an image can be determined from its caption, we demonstrate localization of these concepts upon learning from a statistical affect model for world concepts. The affect model is derived from non-invasively acquired fixation patterns on labeled images, and guides localization of affective objects (faces, reptiles) and actions (look, read) from fixations in unlabeled images. Experimental results obtained on a database of 500 images confirm the effectiveness and promise of the proposed approach.
Subramanian Ramanathan, Harish Katti, Raymond Huang, Tat-Seng Chua, Mohan Kankanhalli
ACM Multimedia1
2008 Impact of vertex clustering on registration-based 3D dynamic mesh coding
Subramanian Ramanathan, Ashraf A. Kassim, Tiow Seng Tan
Image Vis. Comput.1
2006 Human Facial Expression Recognition using a 3D Morphable Model
abstract
We propose a novel approach to the detection and classification of human facial expressions using a morphable 3D model. We acquire the various expressions of an individual using a face scanner that produces textured 3D meshes using stereoscopic reconstruction. A morphable expression model (MEM), that incorporates emotion-dependent face variations in terms of morphing parameters, is then computed by establishing correspondence among the emotive faces. These morphing parameters are used for emotion recognition and classification. We demonstrate that the different facial expressions correspond to distinct clusters in the expression space.
Subramanian Ramanathan, Ashraf A. Kassim, Y. V. Venkatesh, Wu Sin Wah
ICIP1
2004 Multi-resolution streaming and rendering of 3-D dynamic data
Subramanian Ramanathan, Ashraf A. Kassim, Kuntal Sengupta
ICIP1
1993 Scheduling algorithms for multihop radio networks
abstract
Algorithms for transmission scheduling in multihop broadcast radio networks are presented. Both link scheduling and broadcast scheduling are considered. In each instance, scheduling algorithms are given that improve upon existing algorithms both theoretically and experimentally. It is shown that tree networks can be scheduled optimally and that arbitrary networks can be scheduled so that the schedule is bounded by a length that is proportional to a function of the network thickness times the optimum. Previous algorithms could guarantee only that the schedules were bounded by a length no worse than the maximum node degree times optimum. Since the thickness is typically several orders of magnitude less than the maximum node degree, the algorithms presented represent a considerable theoretical improvement. Experimentally, a realistic model of a radio network is given and the performance of the new algorithms is studied. These results show that, for both types of scheduling, the new algorithms (experimentally) perform consistently better than earlier methods.>
Subramanian Ramanathan, Errol L. Lloyd
IEEE/ACM Trans. Netw.1
1992 Scheduling Algorithms for Multi-Hop Radio Networks
abstract
New algorithms for transmission scheduling in multihop broadcast radio networks are presented. Both link scheduling and broadcast scheduling are considered. In each instance scheduling algorithms are given that improve upon existing algorithms both theoretically and experimentally. Theoretically, it is shown that tree networks can be scheduled optimally, and that arbitrary networks can be scheduled so that the schedule is bounded by a length that is proportional to a function of the network thickness times the optimum. Previous algorithms could guarantee only that the schedules were bounded by a length no worse than the maximum node degree, the algorithms presented here represent a considerable theoretical improvement. Experimentally, a realistic model of a radio network is given and the performance of the new algorithms is studied. These results show that, for both types of scheduling, the new algorithms (experimentally) perform consistently better than earlier methods.
Subramanian Ramanathan, Errol L. Lloyd
SIGCOMM1