Krishna Somandepalli

dblp:192/5601 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
16since 2021 · last 2024
0000-0002-2845-1079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 VideoPoet: A Large Language Model for Zero-Shot Video Generation
abstract
We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004
ICML11
2024 A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
abstract
Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we propose a novel training approach to effectively learn arbitrary conditional distributions in the audiovisual space. Our key contribution lies in how we parameterize the diffusion timestep in the forward diffusion process. Instead of the standard fixed diffusion timestep, we propose applying variable diffusion timesteps across the temporal dimension and across modalities of the inputs. This formulation offers flexibility to introduce variable noise levels for various portions of the input, hence the term mixture of noise levels. We propose a transformer-based audiovisual latent diffusion model and show that it can be trained in a task-agnostic fashion using our approach to enable a variety of audiovisual generation tasks at inference time. Experiments demonstrate the versatility of our method in tackling cross-modal and multimodal interpolation tasks in the audiovisual space. Notably, our proposed approach surpasses baselines in generating temporally and perceptually consistent samples conditioned on the input. Project page: neurips13025.github.io
Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang 0004, Aren Jansen, Jacob Walker, Krishna Somandepalli
NeurIPS11
2023 Contextually-Rich Human Affect Perception Using Multimodal Scene Information
abstract
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived by humans rely on multiple contextual cues including social settings, foreground interactions, and ambient visual scenes. In this work, we leverage pretrained vision-language (VLN) models to extract descriptions of foreground context from images. Further, we propose a multimodal context fusion (MCF) module to combine foreground cues with the visual scene and person-based contextual information for emotion prediction. We show the effectiveness of our proposed modular design on two datasets associated with natural scenes and TV shows.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP3
2023 A Dataset for Audio-Visual Sound Event Detection in Movies
abstract
Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at scale. Movies depict various real-life and fictional scenarios which makes them a rich resource for mining a wide range of audio events. In this work, we present a dataset of audio events called Subtitle-Aligned Movie Sounds (SAM-S). We use publicly available closed-caption transcripts to automatically mine over 110K audio events from 430 movies. We identify three dimensions to categorize audio events: sound, source, quality, and present the steps involved to produce a final taxonomy of 245 sounds. We discuss the choices involved in generating the taxonomy, and also highlight the human-centered nature of sounds in our dataset. We establish a baseline performance for audio-only sound classification of 34.76% mean average precision, and show that incorporating visual information can further improve the performance by 5%. Data and code are made available for research at https://github.com/usc-sail/mica-subtitle-aligned-movie-sounds
Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shri Narayanan
ICASSP3
2023 Heterogeneous Graph Learning for Acoustic Event Classification
abstract
Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, graph structure does not appear naturally in audiovisual data. Graphs for audiovisual data are constructed manually which is both difficult and sub-optimal. In this work, we address this problem by (i) proposing a parametric graph construction strategy for the intra-modal edges, and (ii) learning the crossmodal edges. To this end, we develop a new model, heterogeneous graph crossmodal network (HGCN) that learns the crossmodal edges. Our proposed model can adapt to various spatial and temporal scales owing to its parametric construction, while the learnable crossmodal edges effectively connect the relevant nodes across modalities. Experiments on a large benchmark dataset (AudioSet) show that our model is state-of-the-art (0.53 mean average precision), outperforming transformer-based models and other graph-based models. Our code is available at github.com/AmirSh15/Crossmodalitygraph
Amir Shirian, Mona Ahmadian, Krishna Somandepalli, Tanaya Guha
ICASSP3
2023 LanSER: Language-Model Supported Speech Emotion Recognition
abstract
Speech emotion recognition (SER) models typically rely on costly human-labeled data for training, making scaling methods to large speech datasets and nuanced emotion taxonomies difficult. We present LanSER, a method that enables the use of unlabeled data by inferring weak emotion labels via pre-trained large language models through weakly-supervised learning. For inferring weak labels constrained to a taxonomy, we use a textual entailment approach that selects an emotion label with the highest entailment score for a speech transcript extracted via automatic speech recognition. Our experimental results show that models pre-trained on large datasets with this weak supervision outperform other baseline models on standard SER datasets when fine-tuned, and show improved label efficiency. Despite being pre-trained on labels derived only from text, we show that the resulting representations appear to model the prosodic content of speech.
Taesik Gong, Josh Belanich, Krishna Somandepalli, Arsha Nagrani, Brian Eoff, Brendan Jou
INTERSPEECH3
2023 MM-AU: Towards Multimodal Understanding of Advertisement Videos
abstract
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches.
Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan
ACM Multimedia4
2023 MovieCLIP: Visual Scene Recognition in Movies
abstract
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don’t consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shri Narayanan
WACV3
2023 A study of bias mitigation strategies for speaker recognition
Raghuveer Peri, Krishna Somandepalli, Shri Narayanan
Comput. Speech Lang.2
2023 Cross Modal Video Representations for Weakly Supervised Active Speaker Localization
abstract
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is talking, and not. Speaker activity can be automatically discerned from the rich multimodal information present in the media content. This is however a challenging problem due to the vast variety and contextual variability in media content, and the lack of labeled data. In this work, we present a cross-modal neural network for learning visual representations, which have implicit information pertaining to the spatial location of a speaker in the visual frames. Avoiding the need for manual annotations for active speakers in visual frames, acquiring of which is very expensive, we present a weakly supervised system for the task of localizing active speakers in movie content. We use the learned cross-modal visual representations, and provide weak supervision from movie subtitles acting as a proxy for voice activity, thus requiring no manual annotations. Furthermore, we propose an audio-assisted post-processing formulation for the task of active speaker detection. We evaluate the performance of the proposed system on three benchmark datasets: i) AVA active speaker dataset, ii) Visual person clustering dataset, and iii) Columbia datset, and demonstrate the effectiveness of the cross-modal embeddings for localizing active speakers in comparison to fully supervised systems.
Krishna Somandepalli, Shri Narayanan
IEEE Trans. Multim.2
2022 Federated Learning for Affective Computing Tasks
abstract
Federated learning mitigates the need to store user data in a central datastore for machine learning tasks, and is particularly beneficial when working with sensitive user data or tasks. Although successfully used for applications such as improving keyboard query suggestions, it is not studied systematically for modeling affective computing tasks which are often laden with subjective labels and high variability across individuals/raters or even by the same participant. In this paper, we study the federated averaging algorithm FedAvg to model self-reported emotional experience and perception labels on a variety of speech, video and text datasets. We identify two learning paradigms that commonly arise in affective computing tasks: modeling of self-reports (user-as-client), and modeling perceptual judgments such as labeling sentiment of online comments (rater-as-client). In the user-as-client setting, we show that FedAvg generally performs on-par with a non-federated model in classifying self-reports. In the rater-as-client setting, FedAvg consistently performed poorer than its non-federated counterpart. We found that the performance of FedAvg degraded for classes where the inter-rater agreement was moderate to low. To address this finding, we propose an algorithm FedRater that learns client-specific label distributions in federated settings. Our experimental results show that FedRater not only improves the overall classification performance compared to FedAvg but also provides insights for estimating proxies of inter-rater agreement in distributed settings.
Krishna Somandepalli, Brian Eoff, Alan Cowen, Kartik Audhkhasi, Josh Belanich, Brendan Jou
ACII1
2022 Visually-aware Acoustic Event Detection using Heterogeneous Graphs
abstract
Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the joint information. In contrast, we employ heterogeneous graphs to explicitly capture the spatial and temporal relationships between the modalities and represent detailed information about the underlying signal. Using heterogeneous graph approaches to address the task of visually-aware acoustic event classification, which serves as a compact, efficient and scalable way to represent data in the form of graphs. Through heterogeneous graphs, we show efficiently modelling of intra- and inter-modality relationships both at spatial and temporal scales. Our model can easily be adapted to different scales of events through relevant hyperparameters. Experiments on AudioSet, a large benchmark, shows that our model achieves state-of-the-art performance. Our code is available at github.com/AmirSh15/VAED HeterGraph.
Amir Shirian, Krishna Somandepalli, Victor Sanchez, Tanaya Guha
INTERSPEECH2
2022 Robust Character Labeling in Movie Videos: Data Resources and Self-Supervised Feature Adaptation
abstract
Robust face clustering is a vital step in enabling computational understanding of visual character portrayal in media. Face clustering for long-form content is challenging because of variations in appearance and lack of supporting large-scale labeled data. Our work in this paper focuses on two key aspects of this problem: the lack of domain-specific training or benchmark datasets, and adapting face embeddings learned on web images to long-form content, specifically movies. First, we present a dataset of over 169000 face tracks curated from 240 Hollywood movies with weak labels on whether a pair of face tracks belong to the same or a different character. We propose an offline algorithm based on nearest-neighbor search in the embedding space to mine hard-examples from these tracks. We then investigate triplet-loss and multiview correlation-based methods for adapting face embeddings to hard-examples. Our experimental results highlight the usefulness of weakly labeled data for domain-specific feature adaptation. Overall, we find that multiview correlation-based adaptation yields more discriminative and robust face embeddings. Its performance on downstream face verification and clustering tasks is comparable to that of the state-of-the-art results in this domain. We also present the SAIL-Movie Character Benchmark corpus developed to augment existing benchmarks. It consists of racially diverse actors and provides face-quality labels for subsequent error analysis. We hope that the large-scale datasets developed in this work can further advance automatic character labeling in videos. All resources are available freely athttps://sail.usc.edu/~ccmi/multiface.
Krishna Somandepalli, Rajat Hebbar, Shri Narayanan
IEEE Trans. Multim.1
2021 A Computational Tool to Study Vocal Participation of Women in UN-ITU Meetings
abstract
International organizations such as the United Nations drive policies that impact our everyday lives. Diverse representation of people and ideas in the decision making process of such bodies is critical to ensure that the policies work for everyone. One aspect of the representation is the partipants' expressed gender. In this work, we focus on analyzing meetings at the International Telecommunication Union (ITU). These meetings include a moderator who mediates the proceedings between delegates from across the world speaking in different languages. For the purpose of quantifying the participation of delegates, we propose a scalable, human-in-the-loop system to first identify the moderator's speech and estimate the speaking time with respect to gender for all the speakers. Our proposed system includes three main audio modules: speech activity detection, gender identification and moderator verification using a human-labelled speech probe. We then estimate percentage of speaking time controlled for the moderator's speech. We present detailed and multilingual performance evaluation of the component systems using state-of-the-art technologies for these tasks. Finally, we examine the vocal participation of female delegates in the 2018 ITU Plenipotentiary Conference spanning for 18 days and about 108 hours of audio recordings.
Rajat Hebbar, Krishna Somandepalli, Raghuveer Peri, Ruchir Travadi, Tracy Tuplin, Fernando Rivera, Shri Narayanan
CBMI2
2021 Loss Function Approaches for Multi-label Music Tagging
abstract
Given the ever-increasing volume of music created and released every day, it has never been more important to study automatic music tagging. In this paper, we present an ensemble-based convolutional neural network (CNN) model trained using various loss functions for tagging musical genres from audio. We investigate the effect of different loss functions and resampling strategies on prediction performance, finding that using focal loss improves overall performance on the the MTG-Jamendo dataset: an imbalanced, multi-label dataset with over 18,000 songs in the public domain, containing 57 labels. Additionally, we report results from varying the receptive field on our base classifier-a CNN-based architecture trained using Mel spectrograms-which also results in a model performance boost and state-of-the-art performance on the Jamendo dataset. We conclude that the choice of the loss function is paramount for improving on existing methods in music tagging, particularly in the presence of class imbalance.
Dillon Knox, Timothy Greer, Benjamin Ma, Emily Kuo, Krishna Somandepalli, Shri Narayanan
CBMI5
2021 Computational Media Intelligence: Human-Centered Machine Analysis of Media
abstract
Media is created by humans for humans to tell stories. There exists a natural and imminent need for creating human-centered media analytics to illuminate the stories being told and to understand their impact on individuals and society at large. An objective understanding of media content has numerous applications for different stakeholders, from creators to decision-/policy-makers to consumers. Advances in multimodal signal processing and machine learning (ML) can enable detailed and nuanced characterization of media content (of who, what, how, where, and why) at scale. They can also aid our understanding of the impact of media on a range of issues, including individual experiences, behavioral, cultural, and societal trends, and commercial outcomes. Modern deep learning models combined with audiovisual signal processing can analyze entertainment media, such as Film & TV content to quantify gender, age, and race representations. This creates awareness in an objective way that was hitherto impossible. On the other hand, text mining and natural language processing allow nuanced understanding of language use and spoken interactions in media, such as News to track patterns and trends across different contexts. Moreover, advances in human sensing have enabled us to directly measure the influence of media on an individual’s physiology (and brain), while social media analysis enables tracking the societal impact of media content on different cross sections of the society. This article reviews representative methodologies and algorithms, tools, and systems advancing human-centered media understanding through ML in the pursuit of developing computational media intelligence.
Krishna Somandepalli, Tanaya Guha, Victor R. Martinez, Naveen Kumar 0004, Hartwig Adam, Shri Narayanan
Proc. IEEE1
2020 Joint Estimation and Analysis of Risk Behavior Ratings in Movie Scripts
abstract
Exposure to violent, sexual, or substanceabuse content in media increases the willingness of children and adolescents to imitate similar behaviors.Computational methods that identify portrayals of risk behaviors from audio-visual cues are limited in their applicability to films in post-production, where modifications might be prohibitively expensive.To address this limitation, we propose a model that estimates content ratings based on the language use in movie scripts, making our solution available at the earlier stages of creative production.Our model significantly improves the state-of-the-art by adapting novel techniques to learn better movie representations from the semantic and sentiment aspects of a character's language use, and by leveraging the co-occurrence of risk behaviors, following a multi-task approach.Additionally, we show how this approach can be useful to learn novel insights on the joint portrayal of these behaviors, and on the subtleties that filmmakers may otherwise not pick up on.
Victor R. Martinez, Krishna Somandepalli, Yalda T. Uhls, Shri Narayanan
EMNLP (1)2
2020 Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context
abstract
Due to its ability to visualize and measure the dynamics of vocal tract shaping during speech production, real-time magnetic resonance imaging (rtMRI) has emerged as one of the prominent research tools. The ability to track different articulators such as the tongue, lips, velum, and the pharynx is a crucial step toward automating further scientific and clinical analysis. Recently, various researchers have addressed the problem of detecting articulatory boundaries, but those are primarily limited to static-image based methods. In this work, we propose to use information from temporal dynamics together with the spatial structure to detect the articulatory boundaries in rtMRI videos. We train a convolutional LSTM network to detect and label the articulatory contours. We compare the produced contours against reference labels generated by iteratively fitting a manually created subject-specific template. We observe that the proposed method outperforms solely image-based methods, especially for the difficult-to-track articulators involved in airway constriction formation during speech.
S. Ashwin Hebbar, Krishna Somandepalli, Asterios Toutios, Shri Narayanan
ICASSP3
2020 Robust Speaker Recognition Using Unsupervised Adversarial Invariance
abstract
In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker embeddings extracted using a pretrained model onto two lower dimensional embedding spaces. The embedding spaces are learnt to disentangle speaker-discriminative information from all other information present in the audio recordings, without supervision about the acoustic conditions. We analyze the robustness of the proposed embeddings to various sources of variability present in the signal for speaker verification and unsupervised clustering tasks on a large-scale speaker recognition corpus. Our analyses show that the proposed system substantially outperforms the baseline in a variety of challenging acoustic scenarios. Furthermore, for the task of speaker diarization on a real-world meeting corpus, our system shows a relative improvement of 36% in the diarization error rate compared to the state-of-the-art baseline.
Raghuveer Peri, Monisankha Pal, Arindam Jati, Krishna Somandepalli, Shri Narayanan
ICASSP4
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia8
2019 Violence Rating Prediction from Movie Scripts
abstract
Violent content in movies can influence viewers’ perception of the society. For example, frequent depictions of certain demographics as perpetrators or victims of abuse can shape stereotyped attitudes. In this work, we propose to characterize aspects of violent content in movies solely from the language used in the scripts. This makes our method applicable to a movie in the earlier stages of content creation even before it is produced. This is complementary to previous works which rely on audio or video post production. Our approach is based on a broad range of features designed to capture lexical, semantic, sentiment and abusive language characteristics. We use these features to learn a vector representation for (1) complete movie, and (2) for an act in the movie. The former representation is used to train a movie-level classification model, and the latter, to train deep-learning sequence classifiers that make use of context. We tested our models on a dataset of 732 Hollywood scripts annotated by experts for violent content. Our performance evaluation suggests that linguistic features are a good indicator for violent content. Furthermore, our ablation studies show that semantic and sentiment features are the most important predictors of violence in this data. To date, we are the first to show the language used in movie scripts is a strong indicator of violent content. This offers novel computational tools to assist in creating awareness of storytelling.
Victor R. Martinez, Krishna Somandepalli, Karan Singla, Anil Ramakrishna, Yalda T. Uhls, Shri Narayanan
AAAI2
2019 Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation
abstract
Speech activity detection in highly variable acoustic conditions is a challenging task. Many approaches to detect speech activity in such conditions involve an inherent knowledge of the noise types involved. Movie audio can offer an excellent research test-bed for developing speech activity models. A robust speech detection in movie audio is also a crucial step for subsequent content analyses such as audio diarization. Obtaining labels for supervision of such data can be very expensive, and may not be scalable. In this paper, we employ a simple, yet effective approach to obtain speech labels for movie data by coarse aligning the subtitles with movie audio. We compiled a dataset, called Subtitle-aligned Movie Corpus (SAM) of nearly 23 hours of data labelled as speech from ninety-five Hollywood movies. We propose convolutional neural network architectures that use log-mel spectrograms as input features to predict speech at a segment-level, as opposed to frame-level. We show that our models trained on SAM outperform existing baselines on two independent, publicly released movie speech datasets. We have made the SAM corpus and pretrained models publicly available for further research.
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP2
2019 Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders
abstract
Audio-signal acquisition as part of wearable sensing adds an important dimension for applications such as understanding human behaviors. As part of a large study on work place behaviours, we collected audio data from individual hospital staff using custom wearable recorders. The audio features collected were limited to preserve privacy of the interactions in the hospital. A first step towards audio processing is to identify the foreground speech of the person wearing the audio badge. This task is challenging because of the multi-party nature of possible ambulatory interactions, lack of access to speaker information and varying channel and ambient conditions. In this paper, we present a speaker-agnostic approach to foreground detection. We propose a convolutional neural network model to predict foreground regions using a limited set of audio features. We show that these models generalize across the proxy corpora we collected in house to approximately match the deployment environment. The proxy corpora contained full audio and was used as a test-bed to analyze our models in greater detail. We also evaluated the models in the workplace setting to measure speech activity. Our experimental results show promising direction for analyzing workplace behaviors with privacy protected sensing.
Amrutha Nadarajan, Krishna Somandepalli, Shri Narayanan
ICASSP2
2019 Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies
abstract
The ability to robustly cluster faces in movies is a necessary step in understanding media content representations of people along dimensions such as gender and age. Building upon the successes of sparse subspace clustering (SSC) in uncovering the underlying structure of the data, in this paper we propose an algorithm called Constraint Propagation Sparse Subspace Clustering (CP-SSC) for applications such as face clustering in videos where pairwise sample constraints (must-link and cannot-link sample pairs) are available in the processing pipeline since detected faces can be tracked locally in time. We learn the subspace structure while simultaneously incorporating the pairwise constraints to construct a similarity matrix needed for clustering. Our joint formulation uses low-rank matrix completion to propagate the initial pairwise constraints, that are used to reinforce the subspace representation during optimization. We evaluate CP-SSC for clustering faces in movies with pre-trained neural network embeddings as features. We first analyze CP-SSC with synthetic data and then show that it can be effectively used to cluster faces in movie videos. We evaluate our method for two movies annotated in-house and two benchmark movies released publicly. We also compare the performance of our algorithm with other clustering approaches that use pairwise constraint information.
Krishna Somandepalli, Shri Narayanan
ICASSP1
2019 Toward Visual Voice Activity Detection for Unconstrained Videos
abstract
The prevalent audio-based Voice Activity Detection (VAD) systems are challenged by the presence of ambient noise and are sensitive to variations in the type of the noise. The use of information from the visual modality, when available, can help overcome some of the problems of audio-based VAD. Existing visual-VAD systems however do not operate directly on the whole image but require intermediate face detection, face landmark detection and subsequent facial feature extraction from the lip region. In this work we present an end-to-end trainable Hierarchical Context Aware (HiCA) architecture for visual-VAD for videos obtained in unconstrained environments which can be trained with videos as input and audio speech labels as output. The network is designed to account for local and global temporal information in a video sequence. In contrast to existing visual-VAD systems our proposed approach does not rely on face detection and subsequent facial feature extraction. It can obtain a VAD accuracy of 66% on a dataset of Hollywood movie videos just with visual information. Further analysis of the representations learned from our visual-VAD system shows that the network learns to localize on human faces, and sometimes speaking human faces specifically. Our quantitative analysis of the effectiveness of face localization shows that our system performs better than sound-localization networks designed for unconstrained videos.
Krishna Somandepalli, Shri Narayanan
ICIP2
2019 Identifying Therapist and Client Personae for Therapeutic Alliance Estimation
abstract
. We measure the strength of the relation between personae and alliance in two experiments. Our results show that (1) alliance can be explained by the interactions between the discovered character types, and (2) models trained on therapist and client personae achieve significant performance gains compared to competitive supervised baselines. Finally, exploratory analysis reveals important character traits that lead to an improved perception of alliance.
Victor R. Martinez, Nikolaos Flemotomos, Victor Ardulov, Krishna Somandepalli, Simon B. Goldberg, Zac E. Imel, David C. Atkins, Shri Narayanan
INTERSPEECH4
2019 Multiview Shared Subspace Learning Across Speakers and Speech Commands
Krishna Somandepalli, Naveen Kumar 0004, Arindam Jati, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2018 Multimodal Representation of Advertisements Using Segment-level Autoencoders
abstract
Automatic analysis of advertisements (ads) poses an interesting problem for learning multimodal representations. A promising direction of research is the development of deep neural network autoencoders to obtain inter-modal and intra-modal representations. In this work, we propose a system to obtain segment-level unimodal and joint representations. These features are concatenated, and then averaged across the duration of an ad to obtain a single multimodal representation. The autoencoders are trained using segments generated by time-aligning frames between the audio and video modalities with forward and backward context. In order to assess the multimodal representations, we consider the tasks of classifying an ad as funny or exciting in a publicly available dataset of 2,720 ads. For this purpose we train the segment-level autoencoders on a larger, unlabeled dataset of 9,740 ads, agnostic of the test set. Our experiments show that: 1) the multimodal representations outperform joint and unimodal representations, 2) the different representations we learn are complementary to each other, and 3) the segment-level multimodal representations perform better than classical autoencoders and cross-modal representations -- within the context of the two classification tasks. We obtain an improvement of about 5% in classification accuracy compared to a competitive baseline.
Krishna Somandepalli, Victor R. Martinez, Naveen Kumar 0004, Shri Narayanan
ICMI1
2018 Improving Gender Identification in Movie Audio Using Cross-Domain Data
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
INTERSPEECH2
2018 Unsupervised Discovery of Character Dictionaries in Animation Movies
abstract
Automatic content analysis of animation movies can enable an objective understanding of character (actor) representations and their portrayals. It can also help illuminate potential markers of unconscious biases and their impact. However, multimedia analysis of movie content has predominantly focused on live-action features. A dearth of multimedia research in this field is because of the complexity and heterogeneity in the design of animated characters-an extremely challenging problem to be generalized by a single method or model. In this paper, we address the problem of automatically discovering characters in animation movies as a first step toward automatic character labeling in these media. Movie-specific character dictionaries can act as a powerful first step for subsequent content analysis at scale. We propose an unsupervised approach which requires no prior information about the characters in a movie. We first use a deep neural network-based object detector that is trained on natural images to identify a set of initial character candidates. These candidates are further pruned using saliency constraints and visual object tracking. A character dictionary per movie is then generated from exemplars obtained by clustering these candidates. We are able to identify both anthropomorphic and nonanthropomorphic characters in a dataset of 46 animation movies with varying composition and character design. Our results indicate high precision and recall of the automatically detected characters compared to human-annotated ground truth, demonstrating the generalizability of our approach.
Krishna Somandepalli, Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
IEEE Trans. Multim.1
2017 Semantic Edge Detection for Tracking Vocal Tract Air-Tissue Boundaries in Real-Time Magnetic Resonance Images
Krishna Somandepalli, Asterios Toutios, Shri Narayanan
INTERSPEECH1
2016 Articulatory Synthesis Based on Real-Time Magnetic Resonance Imaging Data
Asterios Toutios, Tanner Sorensen, Krishna Somandepalli, Rachel Alexander, Shri Narayanan
INTERSPEECH3