Rajat Hebbar

dblp:226/1866 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0002-0904-0573ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition
abstract
Recent studies have explored using pre-trained embeddings for speech emotion recognition, achieving comparable performance to conventional methods that rely on low-level knowledge-inspired acoustic features. These embeddings are often generated from models trained on large-scale speech datasets using self-supervised or weakly-supervised learning objectives. Despite the significant advancements made in SER through pre-trained embeddings, there is a limited understanding of the trustworthiness of these methods, including privacy breaches, unfair performance, vulnerability to adversarial attacks, and computational cost, all of which may hinder the real-world deployment of these systems. In response, we introduce TrustSER, a general framework designed to evaluate the trustworthiness of SER systems using deep learning methods, focusing on privacy, safety, fairness, and sustainability, offering unique insights into future research in the field of SER. Our code is publicly available under: https://github.com/usc-sail/trust-ser.
Tiantian Feng, Rajat Hebbar, Shri Narayanan
ICASSP2
2024 CVAT-BWV: A Web-Based Video Annotation Platform for Police Body-Worn Video
Parsa Hejabi, Akshay Kiran Padte, Preni Golazizian, Rajat Hebbar, Jackson Trager, Georgios Chochlakis, Aditya Kommineni, Ellie Graeden, Shri Narayanan, Benjamin A. T. Graham, Morteza Dehghani
IJCAI4
2023 Contextually-Rich Human Affect Perception Using Multimodal Scene Information
abstract
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived by humans rely on multiple contextual cues including social settings, foreground interactions, and ambient visual scenes. In this work, we leverage pretrained vision-language (VLN) models to extract descriptions of foreground context from images. Further, we propose a multimodal context fusion (MCF) module to combine foreground cues with the visual scene and person-based contextual information for emotion prediction. We show the effectiveness of our proposed modular design on two datasets associated with natural scenes and TV shows.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP2
2023 A Dataset for Audio-Visual Sound Event Detection in Movies
abstract
Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at scale. Movies depict various real-life and fictional scenarios which makes them a rich resource for mining a wide range of audio events. In this work, we present a dataset of audio events called Subtitle-Aligned Movie Sounds (SAM-S). We use publicly available closed-caption transcripts to automatically mine over 110K audio events from 430 movies. We identify three dimensions to categorize audio events: sound, source, quality, and present the steps involved to produce a final taxonomy of 245 sounds. We discuss the choices involved in generating the taxonomy, and also highlight the human-centered nature of sounds in our dataset. We establish a baseline performance for audio-only sound classification of 34.76% mean average precision, and show that incorporating visual information can further improve the performance by 5%. Data and code are made available for research at https://github.com/usc-sail/mica-subtitle-aligned-movie-sounds
Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shri Narayanan
ICASSP1
2023 Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism
Rimita Lahiri, Tiantian Feng, Rajat Hebbar, Catherine Lord, So Hyun Kim, Shri Narayanan
INTERSPEECH3
2023 Understanding Spoken Language Development of Children with ASD Using Pre-trained Speech Embeddings
Anfeng Xu, Rajat Hebbar, Rimita Lahiri, Tiantian Feng, Lindsay Butler, Lue Shen, Helen Tager-Flusberg, Shri Narayanan
INTERSPEECH2
2023 FedMultimodal: A Benchmark for Multimodal Federated Learning
abstract
Over the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal.
Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan
KDD4
2023 MM-AU: Towards Multimodal Understanding of Advertisement Videos
abstract
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches.
Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan
ACM Multimedia2
2023 SEAR: Semantically-grounded Audio Representations
abstract
Audio supports visual story-telling in movies through the use of different sounds. These sounds are often tied to different visual elements, including foreground entities, the interactions between them as well as background context. Visual captions provide a condensed view of an image, providing a natural language description of entities and the relationships between them. In this work, we utilize visual captions to semantically ground audio representations in a self-supervised setup. We leverage state-of-the-art vision-language models to augment movie datasets with visual captions at scale to the order of 9.6M captions to learn audio representations from over 2500 hours of movie data. We evaluate the utility of the learned representations and show state-of-the art performance on two movie understanding tasks, genre and speaking-style classification, outperforming video based methods and audio baselines. Finally, we show that the learned model can be transferred in a zero-shot manner through application in both movie understanding tasks and general action recognition.
Rajat Hebbar, Digbalay Bose, Shri Narayanan
ACM Multimedia1
2023 MovieCLIP: Visual Scene Recognition in Movies
abstract
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don’t consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shri Narayanan
WACV2
2022 Robust Character Labeling in Movie Videos: Data Resources and Self-Supervised Feature Adaptation
abstract
Robust face clustering is a vital step in enabling computational understanding of visual character portrayal in media. Face clustering for long-form content is challenging because of variations in appearance and lack of supporting large-scale labeled data. Our work in this paper focuses on two key aspects of this problem: the lack of domain-specific training or benchmark datasets, and adapting face embeddings learned on web images to long-form content, specifically movies. First, we present a dataset of over 169000 face tracks curated from 240 Hollywood movies with weak labels on whether a pair of face tracks belong to the same or a different character. We propose an offline algorithm based on nearest-neighbor search in the embedding space to mine hard-examples from these tracks. We then investigate triplet-loss and multiview correlation-based methods for adapting face embeddings to hard-examples. Our experimental results highlight the usefulness of weakly labeled data for domain-specific feature adaptation. Overall, we find that multiview correlation-based adaptation yields more discriminative and robust face embeddings. Its performance on downstream face verification and clustering tasks is comparable to that of the state-of-the-art results in this domain. We also present the SAIL-Movie Character Benchmark corpus developed to augment existing benchmarks. It consists of racially diverse actors and provides face-quality labels for subsequent error analysis. We hope that the large-scale datasets developed in this work can further advance automatic character labeling in videos. All resources are available freely athttps://sail.usc.edu/~ccmi/multiface.
Krishna Somandepalli, Rajat Hebbar, Shri Narayanan
IEEE Trans. Multim.2
2021 A Computational Tool to Study Vocal Participation of Women in UN-ITU Meetings
abstract
International organizations such as the United Nations drive policies that impact our everyday lives. Diverse representation of people and ideas in the decision making process of such bodies is critical to ensure that the policies work for everyone. One aspect of the representation is the partipants' expressed gender. In this work, we focus on analyzing meetings at the International Telecommunication Union (ITU). These meetings include a moderator who mediates the proceedings between delegates from across the world speaking in different languages. For the purpose of quantifying the participation of delegates, we propose a scalable, human-in-the-loop system to first identify the moderator's speech and estimate the speaking time with respect to gender for all the speakers. Our proposed system includes three main audio modules: speech activity detection, gender identification and moderator verification using a human-labelled speech probe. We then estimate percentage of speaking time controlled for the moderator's speech. We present detailed and multilingual performance evaluation of the component systems using state-of-the-art technologies for these tasks. Finally, we examine the vocal participation of female delegates in the 2018 ITU Plenipotentiary Conference spanning for 18 days and about 108 hours of audio recordings.
Rajat Hebbar, Krishna Somandepalli, Raghuveer Peri, Ruchir Travadi, Tracy Tuplin, Fernando Rivera, Shri Narayanan
CBMI1
2019 Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation
abstract
Speech activity detection in highly variable acoustic conditions is a challenging task. Many approaches to detect speech activity in such conditions involve an inherent knowledge of the noise types involved. Movie audio can offer an excellent research test-bed for developing speech activity models. A robust speech detection in movie audio is also a crucial step for subsequent content analyses such as audio diarization. Obtaining labels for supervision of such data can be very expensive, and may not be scalable. In this paper, we employ a simple, yet effective approach to obtain speech labels for movie data by coarse aligning the subtitles with movie audio. We compiled a dataset, called Subtitle-aligned Movie Corpus (SAM) of nearly 23 hours of data labelled as speech from ninety-five Hollywood movies. We propose convolutional neural network architectures that use log-mel spectrograms as input features to predict speech at a segment-level, as opposed to frame-level. We show that our models trained on SAM outperform existing baselines on two independent, publicly released movie speech datasets. We have made the SAM corpus and pretrained models publicly available for further research.
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP1
2018 Improving Gender Identification in Movie Audio Using Cross-Domain Data
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
INTERSPEECH1