Tanaya Guha

dblp:00/8763 · DBLP profile ↗
← Back
60ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0003-2167-4891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Online graph based transforms for intra-predicted imaging data
abstract
Orthogonal transforms are key components of several image and video compression systems and standards, as they provide a de-correlated representation of signals to enhance compressibility. However, the most commonly used transforms for compression, such as the Discrete Cosine transforms (DCT) and Discrete Sine transforms (DST), are fixed and non-adaptive, limiting their ability to capture complex or varying signal characteristics. Graph-based transforms (GBTs) have shown improved energy compaction and reconstruction performance, but face two major limitations: the need to signal graph information in the compressed bitstream, which increases overhead and may complicates decoder synchronization, and a dependency on offline training process, which is highly dependent on the quality and completeness of the training data. To address these issues, this paper introduces a novel framework, GBT-ONL, which learns GBTs online in the context of block-based predictive transform coding. The proposed GBT-ONL framework uses a shallow fully connected neural network to predict the graph Laplacian needed for both the forward and inverse GBT. By relying only on information available during encoding, GBT-ONL eliminates the need to signal additional information in the compressed bitstream, and removes the requirement for any prior offline training. Evaluations on several video sequences show that GBT-ONL outperforms both traditional (non-learnable) transforms and existing learnable transforms in terms of energy compaction, reconstruction error, and compression efficiency, as measured by BD-PSNR and BD-Rate metrics.
Debaleena Roy, Tanaya Guha, Victor Sanchez
Pattern Recognit.2
2025 Boosting Tiny Face Detection in Videos with an Integral Score Framework
abstract
Face detection technology is critical in video surveillance applications. Unfortunately, many of the existing approaches fail to detect tiny faces, especially in videos acquired in uncontrolled public environments. In this work, we address this challenge by proposing a novel detection framework that relies on a new multiscale detector that does not require the input to be re-scaled. Our framework leverages existing detectors to detect relatively large faces along with a new detector specifically designed for tiny faces. The core part of our framework is the use of an integral score to detect tiny faces by densely scanning several support regions at different scales. Although this strategy may pose high computational complexity in high-resolution videos, the integral operations used by our framework significantly reduce the computational cost, making it feasible for high-resolution videos and real-time applications. Experiments on the WIDER FACE and MEVA datasets show a significant improvement in performance, particularly for tiny faces depicted in surveillance videos acquired in uncontrolled environments.
Roberto Leyva, Guodong Shen, Ozan Bahadir, Victor Sanchez, Tanaya Guha
FG5
2025 Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
abstract
A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener’s response to the interlocutor’s speech. Although significant progress has been made in the context of generating co-speech gestures, generating listener’s response has remained a challenge. We introduce the task of generating continuous head motion response of a listener in response to the speaker’s speech in real time. To this end, we propose a graph-based end-to-end crossmodal model that takes interlocutor’s speech audio as input and directly generates head pose angles (roll, pitch, yaw) of the listener in real time. Different from previous work, our approach is completely data-driven, does not require manual annotations or oversimplify head motion to merely nods and shakes. Extensive evaluation on the dyadic interaction sessions on the IEMOCAP dataset shows that our model produces a low overall error (4.5 degrees) and a high frame rate, thereby indicating its deployability in real-world human-robot interaction systems. Our code is available at - https://github. com/bigzen/Active-Listener
Bishal Ghosh, Liying Li 0001, Tanaya Guha
ICASSP3
2025 Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
abstract
For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as a novel task of jointly forecasting a person’s intent to interact with the agent, their attitude towards the agent and the action they will perform, from the agent’s (egocentric) perspective. We propose SocialEgoNet - a graph-based spatiotemporal framework that exploits task dependencies through a hierarchical multitask learning approach. SocialEgoNet uses whole-body skeletons (keypoints from face, hands and body) extracted from only 1 second of video input for high inference speed. For evaluation, we augment an existing egocentric human-agent interaction dataset with new class labels and bounding box annotations. Extensive experiments on this augmented dataset, named JPL-Social, demonstrate real-time inference and superior performance (average accuracy across all tasks: 83.15%) of our model outperforming several competitive baselines. The additional annotations and code are available at github.com/biantongfei/SocialEgoNet.
Tongfei Bian, Yiming Ma 0003, Mathieu Chollet, Victor Sanchez, Tanaya Guha
ICME5
2025 CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
abstract
We propose CLIP-EBC, the first fully CLIP-based model designed for accurate crowd density estimation. Existing classification-based crowd counting frameworks face challenges when directly applied with CLIP. For instance, these methods quantize count values into bordering real-valued bins, which are inconsistent with CLIP’s pretraining corpus. Besides, this quantization strategy also introduces label ambiguity near shared bin borders. To address these issues, we propose the Enhanced Blockwise Classification (EBC) framework, which utilizes integer-valued bins and additionally incorporates a density-map-based loss to further improve count accuracy. Building on this backbone-agnostic framework, CLIP-EBC fully leverages CLIP’s recognition capabilities for crowd counting. Experiments show that EBC improves classification-based methods by up to 44.5% in mean absolute error (MAE) on the UCF-QNRF dataset. Furthermore, CLIP-EBC achieves state-of-the-art performance on the NWPU-Crowd test set, with an MAE of 58.2, surpassing the previous best method by 8.6%. Our code is available at https://github.com/Yiming-M/CLIP-EBC.
Yiming Ma 0003, Victor Sanchez, Tanaya Guha
ICME3
2025 Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust
abstract
Recent advances in AI has enabled automated analysis of complex media content at scale and generate actionable insights regarding character representation along such dimensions as gender and age. Past work focused on quantifying representation from audio/video/text using various ML models, but without having the audience in the loop. We ask, even if character distribution along demographic dimensions are available, how useful are they to the general public? Do they actually trust the numbers generated by AI models? Our work addresses these questions through a user study, while proposing a new AI-based character representation and visualization tool. Our tool based on the Contrastive Language Image Pretraining (CLIP) foundation model to analyze visual screen data to quantify character representation across dimensions of age and gender. We also designed effective visualizations suitable for presenting such analytics to lay audience. Next, we conducted a user study to seek empirical evidence on the usefulness and trustworthiness of the AI-generated results for carefully chosen movies presented in the form of our visualizations. We note that participants were able to understand the analytics from our visualization, and deemed the tool `overall useful'. Participants also indicated a need for more detailed visualizations to include more demographic categories and contextual information of the characters. Participants' trust in AI-based gender and age models is seen to be moderate to low, although they were not against the use of AI in this context.
Evdoxia Taka, Debadyuti Bhattacharya, Joanne Garde-Hansen, Tanaya Guha
ICMI5
2025 Robust Understanding of Human-robot Social Interactions through Multimodal Distillation
abstract
There is a growing need for social robots and intelligent agents that can effectively interact with and support users. For the interactions to be seamless, the agents need to analyse social scenes and behavioural cues from their (robot’s) perspective. Works that model human-agent interactions in social situations are few; and even those existing ones are computationally too intensive to be deployed in real time or perform poorly in real-world scenarios when only limited information is available. We propose a knowledge distillation framework that models social interactions through various multimodal cues, and yet is robust against incomplete and noisy information during inference. We train a teacher model with multimodal input (body, face and hand gestures, gaze, raw images) that transfers knowledge to a student model which relies solely on body pose. Extensive experiments on two publicly available human-robot interaction datasets demonstrate that our student modelachievesan average accuracy gain of 14.75% over competitive baselines on multiple downstream social understanding tasks, even with up to 51% of its input being corrupted. The student model is also highly efficient - less than 1% in size of the teacher model in terms of parameters and its latency is 11.9% of the teacher model.
Tongfei Bian, Mathieu Chollet, Tanaya Guha
ACM Multimedia3
2025 Self-supervised random mask attention GAN in tackling pose-invariant face recognition
abstract
Pose Invariant Face Recognition (PIFR) has significantly advanced with Generative Adversarial Networks (GANs), which rotate face images acquired at any angle to a frontal view for enhanced recognition. However, such frontalization methods typically need ground-truth frontal-view images, often collected under strict laboratory conditions, making it challenging and costly to acquire the necessary training data. Additionally, traditional self-supervised PIFR methods rely on external rendering models for training, further complicating the overall training process. To tackle these two issues, we propose a new framework called Mask Rotate. Our framework introduces a novel training approach that requires no paired ground truth data for the face image frontalization task. Moreover, it eliminates the need for an external rendering model during training. Specifically, our framework simplifies the face image frontalization task by transforming it into a face image completion task. During the inference or testing stage, it employs a reliable pre-trained rendering model to obtain a frontal-view face image, which may have several regions with missing texture due to pose variations and occlusion. Our framework then uses a novel self-supervised Random Mask Attention Generative Adversarial Network (RMAGAN) to fill in these missing regions by considering them as randomly masked regions. Furthermore, our proposed Mask Rotate framework uses a reliable post-processing model designed to improve the visual quality of the face images after frontalization. In comprehensive experiments, the Mask Rotate framework eliminates the requirement for complex computations during training and achieves strong results, both qualitative and quantitative, compared to the state-of-the-art.
Jiashu Liao, Tanaya Guha, Victor Sanchez
Pattern Recognit.2
2024 Assessing Privacy Risks of Attribute Inference Attacks Against Speech-Based Depression Detection System
abstract
Many AI applications now attempt to infer users’ mental health conditions, such as depression, from their speech data. In addition to the spoken words, the speech audio contains information about speaker’s identity and demographic attributes, exposing users to serious privacy risks. Previous efforts have primarily focused on developing deep models that preserve privacy; however, there have been few attempts to systematically assess and quantify privacy risks in such systems. We present the first framework for systematically assessing privacy risks in a multimodal (audio-lexical) depression detection system particularly looking at attribute inference attacks. Unlike past works that considered only white-box gender inference attacks against unimodal systems, our framework designs novel white-box and black-box attacks across multiple modalities against three protected speaker attributes: gender, age and education level. We present extensive results on a large, clinically validated dataset, demonstrating critical vulnerability of depression detection systems, where an adversary can infer speaker attributes with 59% - 68% accuracy even for inputs as short as 10 seconds of speech. Our results offer insights and guidelines to inform the development and benchmarking of privacy-preserving models for speech-based depression detection systems. Our code and data are available at: https://github.com/apr-aia/privacy_risks
Basmah Alsenani, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha
ECAI4
2024 Is Distance a Modality? Multi-Label Learning for Speech-Based Joint Prediction of Attributed Traits and Perceived Distances in 3D Audio Immersive Environments
abstract
To the best of our knowledge, this article presents the first experiments on speech-based Automatic Personality Perception performed in a virtual immersive audio environment. The key-difference compared to all previous works in the literature is that, in a virtual immersive environment, people perceive not only the voice of the speakers, but also their position and distance in space. Therefore, it is possible to investigate for the first time whether people tend to attribute different traits to people speaking at different distances and, if so, whether this makes a difference in terms of Automatic Personality Perception. The experiments were performed over 360 recordings rendered at different distances (120 speakers including 60 female and 60 male). The results show that there are correlations between perceived distance and personality judgments. Furthermore, the experiments show that the performance in Automatic Personality Perception improves when taking perceived distance into account. These results are important because immersive environments are likely to become one of the main technological interfaces through which people interact with one another and with machines.
Evangelia Fringi, Nesreen Alshubaily, Lorenzo Picinali, Stephen A. Brewster, Tanaya Guha, Alessandro Vinciarelli
ICMI5
2024 Detecting in-car VR Motion Sickness from Lower Face Action Units
abstract
This paper presents the first in-car VR motion sickness (VRMS) detection model based on lower face action units (LF-AUs). Initially developed in a simulated in-car environment with 78 participants, the model’s generalizability was later tested in realworld driving conditions. Motion sickness was induced using visual linear motion in the VR headset and physical horizontal rotation via a rotating chair. We used a convolutional neural network (MobileNetV3) to automatically extract LF-AUs from images of the users’ mouth region, captured by the VR headset’s built-in camera. These LF-AUs were then used to train a Support Vector Regression (SVR) model to estimate motion sickness scores. We compared the SVR model’s performance using LF-AUs, pupil diameters, and physiological features (individually and in combination) from the same VR headset. Results showed that both individual LF-AU (right dimple) and combined LF-AUs had significant Pearson correlations with self-reported motion sickness scores and achieved lower root mean squared error compared to pupil diameters. The best detection results were obtained by combining LF-AUs and pupil diameters, while physiological features alone did not yield significant results. The LF-AUs-based model demonstrated encouraging generalizability across different settings in the independent studies.
Gang Li 0011, Tanaya Guha, Ogechi Onuoha, Zhanyan Qiu, Alana Grant, Zejian Feng, Kathariana Pohlmann, Mark McGill, Stephen A. Brewster, Frank E. Pollick
ISMAR2
2024 On the effects of obfuscating speaker attributes in privacy-aware depression detection
Nujud Aloshban, Anna Esposito, Alessandro Vinciarelli, Tanaya Guha
Pattern Recognit. Lett.4
2023 Heterogeneous Graph Learning for Acoustic Event Classification
abstract
Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, graph structure does not appear naturally in audiovisual data. Graphs for audiovisual data are constructed manually which is both difficult and sub-optimal. In this work, we address this problem by (i) proposing a parametric graph construction strategy for the intra-modal edges, and (ii) learning the crossmodal edges. To this end, we develop a new model, heterogeneous graph crossmodal network (HGCN) that learns the crossmodal edges. Our proposed model can adapt to various spatial and temporal scales owing to its parametric construction, while the learnable crossmodal edges effectively connect the relevant nodes across modalities. Experiments on a large benchmark dataset (AudioSet) show that our model is state-of-the-art (0.53 mean average precision), outperforming transformer-based models and other graph-based models. Our code is available at github.com/AmirSh15/Crossmodalitygraph
Amir Shirian, Mona Ahmadian, Krishna Somandepalli, Tanaya Guha
ICASSP4
2023 Privacy Risks in Speech Emotion Recognition: A Systematic Study on Gender Inference Attack
abstract
Increasingly more applications now use deep networks to analyse speaker's affective states. An undesirable side effect is that models trained to perform one task (e.g, emotion from speech) can be attacked to infer other, possibly privacy-sensitive attributes (e.g., gender) of the speaker. The amount of information an attacker can infer through such attacks is called leakage, and this article presents the first systematic study of the interplay between gender leakage and the main characteristics of the attacker model (family, architecture and training condition). To this end, we define various attack scenarios, and perform extensive experiments to analyse privacy risks in Speech Emotion Recognition (SER). Results show that SER models can leak a speaker's gender with an accuracy of 51% to 95% (upper bound) depending on the attack condition. Furthermore, our results provide fresh insights on how to limit the effectiveness of possible attacks and, thereby, to ensure privacy preservation.
Basmah Alsenani, Tanaya Guha, Alessandro Vinciarelli
INTERSPEECH2
2022 Graph-based Transform based on 3D Convolutional Neural Network for Intra-Prediction of Imaging Data
abstract
This paper presents a novel class of Graph-based Transform based on 3D convolutional neural networks (GBT-CNN) within the context of block-based predictive transform coding of imaging data. The proposed GBT-CNN uses a 3D convolutional neural network (3D-CNN) to predict the graph information needed to compute the transform and its inverse, thus reducing the signalling cost to reconstruct the data after transformation. The GBT-CNN outperforms the DCT and DCT /DST, which are commonly employed in current video codecs, in terms of the percentage of energy preserved by a subset of transform coefficients, the mean squared error of the reconstructed data, and the transform coding gain according to evaluations on several video frames and medical images.
Debaleena Roy, Tanaya Guha, Victor Sanchez
DCC2
2022 Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Kyle Min 0001, Sourya Roy, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar
ECCV (35)4
2022 Self-Supervised Frontalization and Rotation Gan with Random Swap for Pose-Invariant Face Recognition
abstract
The task of pose invariant face recognition (PIFR) has recently seen important improvements by introducing generative adversarial networks (GANs). These GAN-based models synthesize a frontal face image from an image in any pose to boost recognition performance. Most of these methods, however, require the ground-truth frontal face image during optimization as they rely on supervised or semi-supervised learning. In this work, we introduce the self-supervised Frontalization and Rotation GAN (FRGAN), which can synthesize a frontal face image from a non-frontal face image. For self-supervision, the synthesized image is rotated back to the original pose based on reconstruction and adversarial losses. To improve performance, the FRGAN uses the Random Swap, a parameter-free data augmentation strategy that swaps key facial regions between the input image and its reconstructed version to force the generator to synthesize more realistic images. Our qualitative and quantitative experiments on benchmark datasets confirm the strong performance of the FRGAN compared to the state-of-the-art (SOTA).
Jiashu Liao, Victor Sanchez, Tanaya Guha
ICIP3
2022 Fusioncount: Efficient Crowd Counting Via Multiscale Feature Fusion
abstract
State-of-the-art crowd counting models follow an encoder-decoder approach. Images are first processed by the encoder to extract features. Then, to account for perspective distortion, the highest-level feature map is fed to extra components to extract multiscale features, which are the input to the decoder to generate crowd densities. However, in these methods, features extracted at earlier stages during encoding are underutilised, and the multiscale modules can only capture a limited range of receptive fields, albeit with considerable computational cost. This paper proposes a novel crowd counting architecture (FusionCount), which exploits the adaptive fusion of a large majority of encoded features instead of relying on additional extraction components to obtain multiscale features. Thus, it can cover a more extensive scope of receptive field sizes and lower the computational cost. We also introduce a new channel reduction block, which can extract saliency information during decoding and further enhance the model’s performance. Experiments on two benchmark databases demonstrate that our model achieves state-of-the-art results with reduced computational complexity. PyTorch implementation of the model and weights trained on these two datasets are available at https://github.com/YimingMa/FusionCount.
Yiming Ma 0003, Victor Sanchez, Tanaya Guha
ICIP3
2022 Visually-aware Acoustic Event Detection using Heterogeneous Graphs
abstract
Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the joint information. In contrast, we employ heterogeneous graphs to explicitly capture the spatial and temporal relationships between the modalities and represent detailed information about the underlying signal. Using heterogeneous graph approaches to address the task of visually-aware acoustic event classification, which serves as a compact, efficient and scalable way to represent data in the form of graphs. Through heterogeneous graphs, we show efficiently modelling of intra- and inter-modality relationships both at spatial and temporal scales. Our model can easily be adapted to different scales of events through relevant hyperparameters. Experiments on AudioSet, a large benchmark, shows that our model achieves state-of-the-art performance. Our code is available at github.com/AmirSh15/VAED HeterGraph.
Amir Shirian, Krishna Somandepalli, Victor Sanchez, Tanaya Guha
INTERSPEECH4
2022 Multi-Camera Trajectory Forecasting With Trajectory Tensors
abstract
We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajectory of a moving object across a network of cameras. While multi-camera setups are widespread for applications such as surveillance and traffic monitoring, existing trajectory forecasting methods typically focus on single-camera trajectory forecasting (SCTF), limiting their use for such applications. Furthermore, using a single camera limits the field-of-view available, making long-term trajectory forecasting impossible. We address these shortcomings of SCTF by developing an MCTF framework that simultaneously uses all estimated relative object locations from several viewpoints and predicts the object's future location in all possible viewpoints. Our framework follows a Which-When-Where approach that predicts in which camera(s) the objects appear and when and where within the camera views they appear. To this end, we propose the concept of trajectory tensors: a new technique to encode trajectories across multiple camera views and the associated uncertainties. We develop several encoder-decoder MCTF models for trajectory tensors and present extensive experiments on our own database (comprising 600 hours of video data from 15 camera views) created particularly for the MCTF task. Results show that our trajectory tensor models outperform coordinate trajectory-based MCTF models and existing SCTF methods adapted for MCTF.
Olly Styles, Tanaya Guha, Victor Sanchez
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Dynamic Emotion Modeling With Learnable Graphs and Graph Inception Network
abstract
Human emotion is expressed, perceived and captured using a variety of dynamic data modalities, such as speech (verbal), videos (facial expressions) and motion sensors (body gestures). We propose a generalized approach to emotion recognition that can adapt across modalities by modeling dynamic data as structured graphs. The motivation behind the graph approach is to build compact models without compromising on performance. To alleviate the problem of optimal graph construction, we cast this as a joint graph learning and classification task. To this end, we present the Learnable Graph Inception Network (L-GrIN) that jointly learns to recognize emotion and to identify the underlying graph structure in the dynamic data. Our architecture comprises multiple novel components: a new graph convolution operation, a graph inception layer, learnable adjacency, and a learnable pooling function that yields a graph-level embedding. We evaluate the proposed architecture on five benchmark emotion recognition databases spanning three different modalities (video, audio, motion capture), where each database captures one of the following emotional cues: facial expressions, speech and body gestures. We achieve state-of-the-art performance on all five databases outperforming several competitive baselines and relevant existing methods. Our graph architecture shows superior performance with significantly fewer parameters (compared to convolutional or recurrent neural networks) promising its applicability to resource-constrained devices. Our code is available athttps://github.com/AmirSh15/graph_emotion_recognition.
Amir Shirian, Subarna Tripathi, Tanaya Guha
IEEE Trans. Multim.3
2021 Graph Based Transforms based on Graph Neural Networks for Predictive Transform Coding
abstract
This paper introduces the GBT-NN, a novel class of Graph-based Transform within the context of block-based predictive transform coding using intra-prediction. The GBT-NNis constructed by learning a mapping function to map a graph Laplacian representing the covariance matrix of the current block. Our objective of learning such a mapping functionis to design a GBT that performs as well as the KLT without requiring to explicitly com-pute the covariance matrix for each residual block to be transformed. To avoid signallingany additional information required to compute the inverse GBT-NN, we also introduce acoding framework that uses a template-based prediction to predict residuals at the decoder. Evaluation results on several video frames and medical images, in terms of the percentageof preserved energy and mean square error, show that the GBT-NN can outperform the DST and DCT.
Debaleena Roy, Tanaya Guha, Victor Sanchez
DCC2
2021 Compact Graph Architecture for Speech Emotion Recognition
abstract
We propose a deep graph approach to address the task of speech emotion recognition. A compact, efficient and scalable way to represent data is in the form of graphs. Following the theory of graph signal processing, we propose to model speech signal as a cycle graph or a line graph. Such graph structure enables us to construct a Graph Convolution Network (GCN)-based architecture that can perform an accurate graph convolution in contrast to the approximate convolution used in standard GCNs. We evaluated the performance of our model for speech emotion recognition on the popular IEMOCAP and MSP-IMPROV databases. Our model outperforms standard GCN and other relevant deep graph architectures indicating the effectiveness of our approach. When compared with existing speech emotion recognition methods, our model achieves comparable performance to the state-of-the-art with significantly fewer learnable parameters (~30K) indicating its applicability in resource-constrained devices. Our code is available at /github.com/AmirSh15/Compact_SER.
Amir Shirian, Tanaya Guha
ICASSP2
2021 In Defense of Scene Graphs for Image Captioning
abstract
The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics, such as object entities, relationships and attributes. Several studies have noted that the naive use of scene graphs from a black-box scene graph generator harms image captioning performance and that scene graph-based captioning models have to incur the overhead of explicit use of image features to generate decent captions. Addressing these challenges, we propose SG2Caps, a framework that utilizes only the scene graph labels for competitive image captioning performance. The basic idea is to close the semantic gap between the two scene graphs - one derived from the input image and the other from its caption. In order to achieve this, we leverage the spatial location of objects and the Human-Object-Interaction (HOI) labels as an additional HOI graph. SG2Caps outperforms existing scene graph-only captioning models by a large margin, indicating scene graphs as a promising representation for image captioning. Direct utilization of scene graph labels avoids expensive graph convolutions over high-dimensional CNN features resulting in 49% fewer trainable parameters. Our code is available at: https://github.com/Kien085/SG2Caps
Kien Nguyen 0006, Subarna Tripathi, Bang Du, Tanaya Guha, Truong Q. Nguyen
ICCV4
2021 Head Matters: Explainable Human-centered Trait Prediction from Head Motion Dynamics
abstract
We demonstrate the utility of elementary head-motion units termed kinemes for behavioral analytics to predict personality and interview traits. Transforming head-motion patterns into a sequence of kinemes facilitates discovery of latent temporal signatures characterizing the targeted traits, thereby enabling both efficient and explainable trait prediction. Utilizing Kinemes and Facial Action Coding System (FACS) features to predict (a) OCEAN personality traits on the First Impressions Candidate Screening videos, and (b) Interview traits on the MIT dataset, we note that: (1) A Long-Short Term Memory (LSTM) network trained with kineme sequences performs better than or similar to a Convolutional Neural Network (CNN) trained with facial images; (2) Accurate predictions and explanations are achieved on combining FACS action units (AUs) with kinemes, and (3) Prediction performance is affected by the time-length over which head and facial movements are observed.
Surbhi Madan, Monika Gahalawat, Tanaya Guha, Subramanian Ramanathan
ICMI3
2021 Computational Media Intelligence: Human-Centered Machine Analysis of Media
abstract
Media is created by humans for humans to tell stories. There exists a natural and imminent need for creating human-centered media analytics to illuminate the stories being told and to understand their impact on individuals and society at large. An objective understanding of media content has numerous applications for different stakeholders, from creators to decision-/policy-makers to consumers. Advances in multimodal signal processing and machine learning (ML) can enable detailed and nuanced characterization of media content (of who, what, how, where, and why) at scale. They can also aid our understanding of the impact of media on a range of issues, including individual experiences, behavioral, cultural, and societal trends, and commercial outcomes. Modern deep learning models combined with audiovisual signal processing can analyze entertainment media, such as Film & TV content to quantify gender, age, and race representations. This creates awareness in an objective way that was hitherto impossible. On the other hand, text mining and natural language processing allow nuanced understanding of language use and spoken interactions in media, such as News to track patterns and trends across different contexts. Moreover, advances in human sensing have enabled us to directly measure the influence of media on an individual’s physiology (and brain), while social media analysis enables tracking the societal impact of media content on different cross sections of the society. This article reviews representative methodologies and algorithms, tools, and systems advancing human-centered media understanding through ML in the pursuit of developing computational media intelligence.
Krishna Somandepalli, Tanaya Guha, Victor R. Martinez, Naveen Kumar 0004, Hartwig Adam, Shri Narayanan
Proc. IEEE2
2020 Variational Recurrent Sequence-to-Sequence Retrieval for Stepwise Illustration
Vishwash Batra, Aparajita Haldar, Yulan He 0001, Hakan Ferhatosmanoglu, George Vogiatzis, Tanaya Guha
ECIR (1)6
2020 Ensemble Network For Ranking Images Based On Visual Appeal
abstract
We propose a computational framework for ranking images (group photos in particular) taken at the same event within a short time span. The ranking is expected to correspond with human perception of overall appeal of the images. We hypothesize and provide evidence through subjective analysis that the factors that appeal to humans are its emotional content, aesthetics and image quality. We propose a network which is an ensemble of three information channels, each predicting a score corresponding to one of three visual appeal factors. For group emotion estimation, we propose a convolutional neural network (CNN) based architecture for predicting group emotion from images. This new architecture enforces the network to put emphasis on the important regions in the images, and achieves comparable results to the state-of-the-art. Next, we develop a network for the image ranking task that combines group emotion, aesthetics and image quality scores. Owing to the unavailability of suitable databases, we created a new database of manually annotated group photos taken during various social events. We present experimental results on this database and other benchmark databases whenever available. Overall, our experiments show that the proposed framework can reliably predict the overall appeal of images with results closely corresponding to human ranking.
Victor Sanchez, Tanaya Guha
ICASSP3
2020 Attention Selective Network For Face Synthesis And Pose-Invariant Face Recognition
abstract
Face recognition algorithms have improved significantly in recent years since the introduction of deep learning and the availability of large training datasets. However, their performance is still inadequate when the face pose varies as pose variation can dramatically increase intra-person variability. This work proposes a novel generative adversarial architecture called the Attention Selective Network (ASN) to address the problem of pose-invariant face recognition. The ASN introduces an efficient attention mechanism and a multi-part loss function to generate realistic-looking frontal face images from other face poses that can be used for recognizing faces under various poses. Thanks to the high quality of the synthesized images, the ASN achieves superior performance in terms of recognition rates compared to the state-of-the-art supervised methods.
Jiashu Liao, Alex Chichung Kot, Tanaya Guha, Victor Sanchez
ICIP3
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia1
2020 Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classification and Retrieval of Videos
abstract
We present an audio-visual multimodal approach for the task of zero-shot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent past but has primarily been limited to visual modality and to images. We demonstrate that both audio and visual modalities are important for ZSL for videos. Since a dataset to study the task is currently not available, we also construct an appropriate multimodal dataset with 33 classes containing 156, 416 videos, from an existing large scale audio event dataset. We empirically show that the performance improves by adding audio modality for both tasks of zero-shot classification and retrieval, when using multi-modal extensions of embedding learning methods. We also propose a novel method to predict the `dominant' modality using a jointly learned modality attention network. We learn the attention in a semi-supervised setting and thus do not require any additional explicit labelling for the modalities. We provide qualitative validation of the modality specific attention, which also successfully generalizes to unseen test classes.
Kranti Kumar Parida, Neeraj Matiyali, Tanaya Guha
WACV3
2020 Multiple Object Forecasting: Predicting Future Object Locations in Diverse Environments
abstract
This paper introduces the problem of multiple object forecasting (MOF), in which the goal is to predict future bounding boxes of tracked objects. In contrast to existing works on object trajectory forecasting which primarily consider the problem from a birds-eye perspective, we formulate the problem from an object-level perspective and call for the prediction of full object bounding boxes, rather than trajectories alone. Towards solving this task, we introduce the Citywalks dataset, which consists of over 200k high-resolution video frames. Citywalks comprises of footage recorded in 21 cities from 10 European countries in a variety of weather conditions and over 3.5k unique pedestrian trajectories. For evaluation, we adapt existing trajectory forecasting methods for MOF and confirm cross-dataset generalizability on the MOT-17 dataset without fine-tuning. Finally, we present STED, a novel encoder-decoder architecture for MOF. STED combines visual and temporal features to model both object-motion and ego-motion, and outperforms existing approaches for MOF. Code & dataset link: https://github.com/olly-styles/Multiple-Object-Forecasting
Olly Styles, Tanaya Guha, Victor Sanchez
WACV2
2020 Dynamic character graph via online face clustering for movie analysis
abstract
Abstract An effective approach to automated movie content analysis involves building a network (graph) of its characters. Existing work usually builds a static character graph to summarize the content using metadata, scripts or manual annotations. We propose an unsupervised approach to building a dynamic character graph that captures the temporal evolution of character interaction. We refer to this as the character interaction graph (CIG). Our approach has two components: (i) an online face clustering algorithm that discovers the characters in the video stream as they appear, and (ii) simultaneous creation of a CIG using the temporal dynamics of the resulting clusters. We demonstrate the usefulness of the CIG for two movie analysis tasks: narrative structure (acts) segmentation and major character retrieval. Our evaluation on full-length movies containing more than 5000 face tracks shows that the proposed approach achieves superior performance for both the tasks.
Prakhar Kulshreshtha, Tanaya Guha
Multim. Tools Appl.2
2019 Graph-Based Transform with Weighted Self-Loops for Predictive Transform Coding Based on Template Matching
abstract
This paper introduces the GBT-L, a novel class of Graph-based Transform within the context of block-based predictive transform coding. The GBT-L is constructed using a 2D graph with unit edge weights and weighted self-loops in every vertex. The weighted selfloops are selected based on the residual values to be transformed. To avoid signalling any additional information required to compute the inverse GBT-L, we also introduce a coding framework that uses a template-based strategy to predict residual blocks in the pixel and residual domains. Evaluation results on several video frames and medical images, in terms of the percentage of preserved energy and mean square error, show that the GBT-L can outperform the DST, DCT and the Graph-based Separable Transform.
Debaleena Roy, Tanaya Guha, Victor Sanchez
DCC2
2019 Computational Analysis of Gaze Behavior in Autism During Interaction with Virtual Agents
abstract
Individuals with Autism spectrum disorder (ASD) are known to have significantly impaired social interaction and communication abilities. These impairments are characterized by their difficulties in using and perceiving non-verbal cues, such as facial expressions. The difficulty in processing communicators facial expressions is often attributed to the atypical gaze patterns in individuals with ASD. We present a computational study of gaze behavior in adolescents with ASD during their interaction with virtual agents (avatars) in a virtual reality based social communication platform. We study the implications on the subjects pupil response (pupil diameter changes) and looking pattern (fixation coordinates and duration) when exposed to the avatars demonstrating context-relevant emotional expressions. The data related to fixation and pupil response is collected using a commercial eye-tracker for subjects with and without ASD during their interactions with the avatars. This data is analyzed to investigate how the pupil response dynamics and fixation patterns of the ASD group differ from their typically developing peers. Our results indicate that communicators facial expressions can significantly affect the gaze behavior of the ASD subjects. We also observe reduced complexity in the pupil response dynamics, and lower synchrony between pupil response and fixation pattern in the ASD group.
Zeeshan Akhtar, Tanaya Guha
ICASSP2
2019 Learning Affective Correspondence between Music and Image
abstract
We introduce the problem of learning affective correspondence between audio (music) and visual data (images). For this task, a music clip and an image are considered similar (having true correspondence) if they have similar emotion content. In order to estimate this crossmodal, emotion-centric similarity, we propose a deep neural network architecture that learns to project the data from the two modalities to a common representation space, and performs a binary classification task of predicting the affective correspondence (true or false). To facilitate the current study, we construct a large scale database containing more than 3, 500 music clips and 85, 000 images with three emotion classes (positive, neutral, negative). The proposed approach achieves 61.67% accuracy for the affective correspondence prediction task on this database, outperforming two relevant and competitive baselines. We also demonstrate that our network learns modality-specific representations of emotion (without explicitly being trained with emotion labels), which are useful for emotion recognition in individual modalities.
Eeshan Gunesh Dhekane, Tanaya Guha
ICASSP3
2019 Dirichlet Latent Variable Model: A Dynamic Model Based on Dirichlet Prior for Audio Processing
abstract
We propose a dynamic latent variable model for learning latent bases from time varying, non-negative data. We take a probabilistic approach to modeling the temporal dependence in data by introducing a dynamic Dirichlet prior—a Dirichlet distribution with dynamic parameters. This new distribution allows us to assure non-negativity and avoid intractability when sequential updates are performed (otherwise encountered in using Dirichlet prior). We refer to the proposed model as the Dirichlet latent variable model (DLVM). We develop an expectation maximization algorithm for the proposed model, and also derive a maximuma posterioriestimate of the parameters. Furthermore, we connect the proposed DLVM to two popular latent basis learning methods—probabilistic latent component analysis (PLCA) and non-negative matrix factorization (NMF). We show that 1) PLCA is a special case of our DLVM, and 2) DLVM can be interpreted as a dynamic version of NMF. The usefulness of DLVM is demonstrated for three audio processing applications—speaker source separation, denoising, and bandwidth expansion. To this end, a new algorithm for source separation is also proposed. Through extensive experiments on benchmark databases, we show that the proposed model outperforms several relevant existing methods in all three applications.
Anurendra Kumar, Tanaya Guha, Prasanta Kumar Ghosh
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 A Dynamic Latent Variable Model for Source Separation
abstract
We propose a novel latent variable model for learning latent bases for time-varying non-negative data. Our model uses a mixture multinomial as the likelihood function and proposes a Dirichlet distribution with dynamic parameters as a prior, which we call the dynamic Dirichlet prior. An expectation maximization (EM) algorithm is developed for estimating the parameters of the proposed model. Furthermore, we connect our proposed dynamic Dirichlet latent variable model (dynamic DLVM) to the two popular latent basis learning methods - probabilistic latent component analysis (PLCA) and non-negative matrix factorization (NMF). We show that (i) PLCA is a special case of the dynamic DLVM, and (ii) dynamic DLVM can be interpreted as a dynamic version of NMF. The effectiveness of the proposed model is demonstrated through extensive experiments on speaker source separation, and speech-noise separation. In both cases, our method performs better than relevant and competitive baselines. For speaker separation, dynamic DLVM shows 1.38 dB improvement in terms of source to interference ratio, and 1 dB improvement in source to artifact ratio.
Anurendra Kumar, Tanaya Guha, Prasanta Ghosh
ICASSP2
2018 An Online Algorithm for Constrained Face Clustering in Videos
abstract
We address the problem of face clustering in long, real world videos. This is a challenging task because faces in such videos exhibit wide variability in scale, pose, illumination, expressions, and may also be partially occluded. The majority of the existing face clustering algorithms are offline, i.e., they assume the availability of the entire data at once. However, in many practical scenarios, complete data may not be available at the same time or may be too large to process or may exhibit significant variation in the data distribution over time. We propose an online clustering algorithm that processes data sequentially in short segments of variable length. The faces detected in each segment are either assigned to an existing cluster or are used to create a new one. Our algorithm uses several spatiotemporal constraints, and a convolutional neural network (CNN) to obtain a robust representation of the faces in order to achieve high clustering accuracy on two benchmark video databases (82.1 % and 93.8%). Despite being an online method (usually known to have lower accuracy), our algorithm achieves comparable or better results than state-of-the-art offline and online methods.
Prakhar Kulshreshtha, Tanaya Guha
ICIP2
2018 Learning Spontaneity to Improve Emotion Recognition in Speech
abstract
We investigate the effect and usefulness of spontaneity (i.e. whether a given speech is spontaneous or not) in speech in the context of emotion recognition. We hypothesize that emotional content in speech is interrelated with its spontaneity, and use spontaneity classification as an auxiliary task to the problem of emotion recognition. We propose two supervised learning settings that utilize spontaneity to improve speech emotion recognition: a hierarchical model that performs spontaneity detection before performing emotion recognition, and a multitask learning model that jointly learns to recognize both spontaneity and emotion. Through various experiments on the well known IEMOCAP database, we show that by using spontaneity detection as an additional task, significant improvement can be achieved over emotion recognition systems that are unaware of spontaneity. We achieve state-of-the-art emotion recognition accuracy (4-class, 69.1%) on the IEMOCAP database outperforming several relevant and competitive baselines.
Karttikeya Mangalam, Tanaya Guha
INTERSPEECH2
2018 Multichannel Attention Network for Analyzing Visual Behavior in Public Speaking
abstract
We investigate the importance of human centered visual cues for predicting the popularity of a public lecture. We construct a large database of more than 1800 TED talk videos and leverage the corresponding (online) viewers' ratings from YouTube for a measure of popularity of the TED talks. Visual cues related to facial and physical appearance, facial expressions, and pose variations are learned using convolutional neural networks (CNN) connected to an attention-based long short-term memory (LSTM) network to predict the video popularity. The proposed overall network is end-to-end-trainable, and achieves state-of-the-art prediction accuracy indicating that the visual cues alone contain highly predictive information about the popularity of a talk. We also demonstrate qualitatively that the network learns a human-like attention mechanism, which is particularly useful for interpretability, i.e. how attention varies with time, and across different visual cues as a function of their relative importance.
Tanaya Guha
WACV2
2018 A Computational Study of Expressive Facial Dynamics in Children with Autism
abstract
Several studies have established that facial expressions of children with autism are often perceived as atypical, awkward or less engaging by typical adult observers. Despite this clear deficit in the quality of facial expression production, very little is understood about its underlying mechanisms and characteristics. This paper takes a computational approach to studying details of facial expressions of children with high functioning autism (HFA). The objective is to uncover those characteristics of facial expressions, notably distinct from those in typically developing children, and which are otherwise difficult to detect by visual inspection. We use motion capture data obtained from subjects with HFA and typically developing subjects while they produced various facial expressions. This data is analyzed to investigate how the overall and local facial dynamics of children with HFA differ from their typically developing peers. Our major observations include reduced complexity in the dynamic facial behavior of the HFA group arising primarily from the eye region.
Tanaya Guha, Ruth B. Grossman, Shri Narayanan
IEEE Trans. Affect. Comput.1
2018 Unsupervised Discovery of Character Dictionaries in Animation Movies
abstract
Automatic content analysis of animation movies can enable an objective understanding of character (actor) representations and their portrayals. It can also help illuminate potential markers of unconscious biases and their impact. However, multimedia analysis of movie content has predominantly focused on live-action features. A dearth of multimedia research in this field is because of the complexity and heterogeneity in the design of animated characters-an extremely challenging problem to be generalized by a single method or model. In this paper, we address the problem of automatically discovering characters in animation movies as a first step toward automatic character labeling in these media. Movie-specific character dictionaries can act as a powerful first step for subsequent content analysis at scale. We propose an unsupervised approach which requires no prior information about the characters in a movie. We first use a deep neural network-based object detector that is trained on natural images to identify a set of initial character candidates. These candidates are further pruned using saliency constraints and visual object tracking. A character dictionary per movie is then generated from exemplars obtained by clustering these candidates. We are able to identify both anthropomorphic and nonanthropomorphic characters in a dataset of 46 animation movies with varying composition and character design. Our results indicate high precision and recall of the automatically detected characters compared to human-annotated ground truth, demonstrating the generalizability of our approach.
Krishna Somandepalli, Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
IEEE Trans. Multim.3
2017 On the role of head motion in affective expression
abstract
Non-verbal behavioral cues, such as head movement, play a significant role in human communication and affective expression. Although facial expression and gestures have been extensively studied in the context of emotion understanding, the head motion (which accompany both) is relatively less understood. This paper studies the significance of head movement in adult's affect communication using videos from movies. These videos are taken from the Acted Facial Expression in the Wild (AFEW) database and are labeled with seven basic emotion categories: anger, disgust, fear, joy, neutral, sadness, and surprise. Considering human head as a rigid body, we estimate the head pose at each video frame in terms of the three Euler angles, and obtain a time-series representation of head motion. First, we investigate the importance of the energy of angular head motion dynamics (displacement, velocity and acceleration) in discriminating among emotions. Next, we analyze the temporal variation of head motion by fitting an autoregressive model to the head motion time series. We observe that head motion carries sufficient information to distinguish any emotion from the rest with high accuracy and this information is complementary to that of facial expression as it helps improve emotion recognition accuracy.
Atanu Samanta, Tanaya Guha
ICASSP2
2017 Music Tempo Estimation Using Sub-Band Synchrony
abstract
Tempo estimation aims at estimating the pace of a musical piece measured in beats per minute. This paper presents a new tempo estimation method that utilizes coherent energy changes across multiple frequency sub-bands to identify the onsets. A new measure, called the sub-band synchrony, is proposed to detect and quantify the coherent amplitude changes across multiple sub-bands. Given a musical piece, our method first detects the onsets using the sub-band synchrony measure. The periodicity of the resulting onset curve, measured using the autocorrelation function, is used to estimate the tempo value. The performance of the sub-band synchrony based tempo estimation method is evaluated on two music databases. Experimental results indicate a reasonable improvement in performance when compared to conventional methods of tempo estimation.
Shreyan Chowdhury, Tanaya Guha, Rajesh M. Hegde
INTERSPEECH2
2016 A multimodal mixture-of-experts model for dynamic emotion prediction in movies
abstract
This paper addresses the problem of continuous emotion prediction in movies from multimodal cues. The rich emotion content in movies is inherently multimodal, where emotion is evoked through both audio (music, speech) and video modalities. To capture such affective information, we put forth a set of audio and video features that includes several novel features such as, Video Compressibility and Histogram of Facial Area (HFA). We propose a Mixture of Experts (MoE)-based fusion model that dynamically combines information from the audio and video modalities for predicting the emotion evoked in movies. A learning module, based on hard Expectation-Maximization (EM) algorithm, is presented for the MoE model. Experiments on a database of popular movies demonstrate that our MoE-based fusion method outperforms popular fusion strategies (e.g. early and late fusion) in the context of dynamic emotion prediction.
Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
ICASSP3
2016 Opening big in box office? Trailer content can help
abstract
Computational prediction of a movie's financial success usually relies only on metadata such as - genre, budget, actors, Motion Picture Association of America (MPAA) rating and critics' reviews. We argue that movie trailers, created to invoke viewers' interest and curiosity about a movie, carry complementary information for predicting a movie's financial future. We created a database consisting of 474 American movie trailers along with various metadata and information about movie's financial success in the opening weekend. A number of features that capture the emotional information contained in the audiovisual stream of the trailers are designed and extracted. We observe that the content-based features have as much predictive information as the meta features. Through regression analysis on our database, we show that signal information from trailer content improves the prediction performance.
Adarsh Tadimari, Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
ICASSP3
2016 A trajectory clustering approach to crowd flow segmentation in videos
abstract
This work proposes a trajectory clustering-based approach for segmenting flow patterns in high density crowd videos. The goal is to produce a pixel-wise segmentation of a video sequence (static camera), where each segment corresponds to a different motion pattern. Unlike previous studies that use only motion vectors, we extract full trajectories so as to capture the complete temporal evolution of each region (block) in a video sequence. The extracted trajectories are dense, complex and often overlapping. A novel clustering algorithm is developed to group these trajectories that takes into account the information about the trajectories' shape, location, and the density of trajectory patterns in a spatial neighborhood. Once the trajectories are clustered, final motion segments are obtained by grouping of the resulting trajectory clusters on the basis of their area of overlap, and average flow direction. The proposed method is validated on a set of crowd videos that are commonly used in this field. On comparison with several state-of-the-art techniques, our method achieves better overall accuracy.
Tanaya Guha
ICIP2
2016 Novel affective features for multiscale prediction of emotion in music
abstract
The majority of computational work on emotion in music concentrates on developing machine learning methodologies to build new, more accurate prediction systems, and usually relies on generic acoustic features. Relatively less effort has been put to the development and analysis of features that are particularly suited for the task. The contribution of this paper is twofold. First, the paper proposes two features that can efficiently capture the emotion-related properties in music. These features are named compressibility and sparse spectral components. These features are designed to capture the overall affective characteristics of music (global features). We demonstrate that they can predict emotional dimensions (arousal and valence) with high accuracy as compared to generic audio features. Secondly, we investigate the relationship between the proposed features and the dynamic variation in the emotion ratings. To this end, we propose a novel Haar transform-based technique to predict dynamic emotion ratings using only global features.
Naveen Kumar 0004, Tanaya Guha, Che-Wei Huang, Colin Vaz, Shri Narayanan
MMSP2
2015 Computationally deconstructing movie narratives: An informatics approach
abstract
In general, popular films and screenplays follow a well defined storytelling paradigm that comprises three essential segments or acts: exposition (act I), conflict (act II) and resolution (act III). Deconstructing a movie into its narrative units can enrich semantic understanding of movies, and help in movie summarization, navigation and detection of the key events. A multimodal framework for detecting such three act narrative structure is developed in this paper. Various low-level features are designed and extracted from video, audio and text channels of a movie so as to capture the pace and excitement of the movie's narrative. Information from the three modalities is combined to compute a continuous dynamic measure of the movie's narrative flow, referred to as the story intensity of the movie in this paper. Guided by the knowledge of film grammar, the act boundaries are detected, and compared against annotations collected from human experts. Promising results are demonstrated for nine full-length Hollywood feature films of various genres.
Tanaya Guha, Naveen Kumar 0004, Shri Narayanan, Stacy L. Smith
ICASSP1
2015 On quantifying facial expression-related atypicality of children with Autism Spectrum Disorder
abstract
by adult observers. This paper focuses on data driven ways to analyze and quantify atypicality in facial expressions of children with ASD. Our objective is to uncover those characteristics of facial gestures that induce the sense of perceived atypicality in observers. Using a carefully collected motion capture database, facial expressions of children with and without ASD are compared within six basic emotion categories employing methods from information theory, time-series modeling and statistical analysis. Our experiments show that children with ASD usually have less complex expression producing mechanisms; the differences in facial dynamics between children with and without ASD primarily come from the eye region. Our study also notes that children with ASD exhibit lower symmetry between left and right regions, and lower variation in motion intensity across facial regions.
Tanaya Guha, Anil Ramakrishna, Ruth B. Grossman, Darren Hedley, Sungbok Lee, Shri Narayanan
ICASSP1
2015 Gender Representation in Cinematic Content: A Multimodal Approach
abstract
The goal of this paper is to enable an objective understanding of gender portrayals in popular films and media through multimodal content analysis. An automated system for analyzing gender representation in terms of screen presence and speaking time is developed. First, we perform independent processing of the video and the audio content to estimate gender distribution of screen presence at shot level, and of speech at utterance level. A measure of the movie's excitement or intensity is computed using audiovisual features for every scene. This measure is used as a weighting function to combine the gender-based screen/speaking time information at shot/utterance level to compute gender representation for the entire movie. Detailed results and analyses are presented on seventeen full length Hollywood movies.
Tanaya Guha, Che-Wei Huang, Naveen Kumar 0004, Shri Narayanan
ICMI1
2014 Learning sparse models for image quality assessment
abstract
Many successful image quality metrics rely on the structural information in an image to assess its perceptual quality. Extracting the structural information that is perceptually meaningful to our visual system, however, is a challenging task. This paper proposes a new quality assessment metric that relies on a sparse modeling approach to learn the inherent structures of the image. These structures are learnt as a set of basis vectors, such that any structure in the image can be represented by a linear combination of only a few of these basis vectors. This strategy is known to generate basis vectors that are qualitatively similar to the receptive field of the simple cells present in the mammalian primary visual cortex. The perceptual quality of the distorted image is estimated by comparing the structures of the reference and the distorted images in terms of the learnt basis vectors. Our approach is evaluated on five standard subject-rated image quality assessment datasets. The proposed metric exhibits high correlation with the subjective ratings outperforming several well established methods.
Tanaya Guha, Ehsan Nezhadarya, Rabab K. Ward
ICASSP1
2014 Sparse representation-based image quality assessment
Tanaya Guha, Ehsan Nezhadarya, Rabab K. Ward
Signal Process. Image Commun.1
2014 Image Similarity Using Sparse Representation and Compression Distance
abstract
A new line of research uses compression methods to measure the similarity between signals. Two signals are considered similar if one can be compressed significantly when the information of the other is known. The existing compression-based similarity methods, although successful in the discrete one dimensional domain, do not work well in the context of images. This paper proposes a sparse representation-based approach to encode the information content of an image using information from the other image, and uses the compactness (sparsity) of the representation as a measure of its compressibility (how much can the image be compressed) with respect to the other image. The sparser the representation of an image, the better it can be compressed and the more it is similar to the other image. The efficacy of the proposed measure is demonstrated through the high accuracies achieved in image clustering, retrieval and classification.
Tanaya Guha, Rabab K. Ward
IEEE Trans. Multim.1
2013 Image similarity measurement from sparse reconstruction errors
abstract
This paper presents a new approach to measuring the similarity between two images using sparse reconstruction. Our approach alleviates the difficulty of selecting and extracting suitable features from images which usually requires domain-specific knowledge. The proposed measure, the Sparse SNR (SSNR), does not use any prior knowledge about the data type or the application. SSNR is generic in the sense that it is applicable, without modification, to a variety of problems involving different types of images. Given a pair of images, a set of basis vectors (dictionary) is learnt for each image such that each image can be represented as a linear combination of a small number of its dictionary elements. Each image is reconstructed by two dictionaries - the one trained on the image itself and the second - trained on the other image. We develop a novel similarity measure based on the resulting reconstruction errors. To the best of our knowledge, this is the first attempt to develop a sparse reconstruction-based similarity measure. Excellent classification, clustering and retrieval results are achieved on benchmark datasets involving facial images and textures.
Tanaya Guha, Rabab K. Ward, Tyseer Aboulnasr
ICASSP1
2012 A sparse reconstruction based algorithm for image and video classification
abstract
The success of sparse reconstruction based classification algorithms largely depends on the choice of overcomplete bases (dictionary). Existing methods either use the training samples as the dictionary elements or learn a dictionary by optimizing a cost function with an additional discriminating component. While the former method requires a good number of training samples per class and is not suitable for video signals, the later adds instability and more computational load. This paper presents a sparse reconstruction based classification algorithm that mitigates the above difficulties. We argue that learning class-specific dictionaries, one per class, is a natural approach to discrimination. We describe each training signal by an error vector consisting of the reconstruction errors the signal produces w.r.t each dictionary. This representation is robust to noise, occlusion and is also highly discriminative. The efficacy of the proposed method is demonstrated in terms of high accuracy for image-based Species and Face recognition and video-based Action recognition.
Tanaya Guha, Rabab K. Ward
ICASSP1
2012 Learning Sparse Representations for Human Action Recognition
abstract
This paper explores the effectiveness of sparse representations obtained by learning a set of overcomplete basis (dictionary) in the context of action recognition in videos. Although this work concentrates on recognizing human movements-physical actions as well as facial expressions-the proposed approach is fairly general and can be used to address other classification problems. In order to model human actions, three overcomplete dictionary learning frameworks are investigated. An overcomplete dictionary is constructed using a set of spatio-temporal descriptors (extracted from the video sequences) in such a way that each descriptor is represented by some linear combination of a small number of dictionary elements. This leads to a more compact and richer representation of the video sequences compared to the existing methods that involve clustering and vector quantization. For each framework, a novel classification algorithm is proposed. Additionally, this work also presents the idea of a new local spatio-temporal feature that is distinctive, scale invariant, and fast to compute. The proposed approach repeatedly achieves state-of-the-art results on several public data sets containing various physical actions and facial expressions.
Tanaya Guha, Rabab K. Ward
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Action recognition by learnt class-specific overcomplete dictionaries
abstract
This paper presents a sparse signal representation based approach to address the problem of human action recognition in videos. For each action, a set of redundant basis (dictionary) is learnt by solving a sparse optimization problem. A dictionary is learnt using the image patches of its corresponding action, such that every patch vector is represented by some linear combination of a small number of basis vectors. By learning one dictionary per action, it is expected that each dictionary can efficiently represent one particular action. We show that such class-specific dictionaries - each representative of one action - provide a powerful means of action classification. Given a query sequence, the classifier seeks the dictionary that best approximates the query class. We have evaluated the proposed approach on the standard datasets. Experimental results demonstrate high accuracy and robustness against occlusion or viewpoint changes.
Tanaya Guha, Rabab K. Ward
FG1
2010 Differential Radon Transform for gait recognition
abstract
Experimental studies have proved that high frequency components have the maximum contribution in silhouette-based gait recognition. The Radon Transform (RT), used in gait analysis for its ability to compute useful directional projections, fails to capture the necessary high frequency content of images. In this paper we present the Differential Radon Transform (DiffRT) - a novel adaptation of the standard RT designed to extract such high frequency information efficiently. The proposed transform is used to extract a set of features from gait silhouettes. We provide both theoretical and experimental evidence that DiffRT can indeed collect the important image information to facilitate gait-based human recognition. Averaged silhouettes from USF database are used for performance evaluation following the gait challenge framework. Our proposed method achieves high recognition accuracy and outperforms several state-of-the-art algorithms.
Tanaya Guha, Rabab K. Ward
ICASSP1