Digbalay Bose

dblp:126/2297 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-5281-1695ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Can Multimodal Foundation Models Help Analyze Child-Inclusive Autism Diagnostic Videos?
Aditya Kommineni, Digbalay Bose, Tiantian Feng, So Hyun Kim, Helen Tager-Flusberg, Somer Bishop, Catherine Lord, Sudarsana Reddy Kadiri, Shri Narayanan
INTERSPEECH2
2024 Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization
abstract
Video summarization remains a huge challenge in computer vision due to the size of the input videos to be summarized. We propose an efficient, language-only video summarizer that achieves competitive accuracy with high data efficiency. Using only textual captions obtained via a zero-shot approach, we train a language transformer model and forego image representations. This method allows us to perform filtration amongst the representative text vectors and condense the sequence. With our approach, we gain explainability with natural language that comes easily for human interpretation and textual summaries of the videos. An ablation study that focuses on modality and data compression shows that leveraging text modality only effectively reduces input data processing while retaining comparable results.
Yoonsoo Nam, Adam Lehavi, Daniel Yang, Digbalay Bose, Swabha Swayamdipta, Shri Narayanan
ICASSP4
2024 Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?
abstract
Multi-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains. Despite its success, the robustness of multi-modal learning for visual recognition is often challenged by the unavailability of a subset of modalities, especially the visual modality. Conventional approaches to mitigate missing modalities in multi-modal learning rely heavily on modality fusion schemes. In contrast, this paper explores the use of text-to-image models to assist multi-modal learning. Specifically, we propose and explore a simple but effective multi-modal learning framework GTI-MM to enhance the data efficiency and model robustness against missing visual modality by imputing the missing data with generative models. Using multiple multi-modal datasets with visual recognition tasks, we present a comprehensive analysis of diverse conditions involving missing visual data. Our findings show that synthetic images benefit training data efficiency with missing visual data during training and improve model robustness with visual data missing during both training and testing. Moreover, we demonstrate GTI-MM is effective with lower generation quantity and simple prompt techniques. Our code base and synthetic images are at https://github.com/usc-sail/GTI-MM.
Tiantian Feng, Daniel Yang, Digbalay Bose, Shri Narayanan
ICMI3
2023 Signal Processing Grand Challenge 2023 - E-Prevention: Sleep Behavior as an Indicator of Relapses in Psychotic Patients
abstract
This paper presents the approach and results of USC SAIL’s submission to the Signal Processing Grand Challenge 2023 – e-Prevention (Task 2), on detecting relapses in psychotic patients. Relapse prediction has proven to be challenging, primarily due to the heterogeneity of symptoms and responses to treatment between individuals. We address these challenges by investigating the use of sleep behavior features to estimate relapse days as outliers in an unsupervised machine learning setting. We extract informative features from human activity and heart rate data collected in the wild, and evaluate various combinations of feature types and time resolutions. We found that short-time sleep behavior features outperformed their awake counterparts and larger time intervals. Our submission was ranked 3rd in the Task’s official leaderboard, demonstrating the potential of such features as an objective and non-invasive predictor of psychotic relapses.
Kleanthis Avramidis, Kranti Adsul, Digbalay Bose, Shri Narayanan
ICASSP3
2023 Contextually-Rich Human Affect Perception Using Multimodal Scene Information
abstract
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived by humans rely on multiple contextual cues including social settings, foreground interactions, and ambient visual scenes. In this work, we leverage pretrained vision-language (VLN) models to extract descriptions of foreground context from images. Further, we propose a multimodal context fusion (MCF) module to combine foreground cues with the visual scene and person-based contextual information for emotion prediction. We show the effectiveness of our proposed modular design on two datasets associated with natural scenes and TV shows.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP1
2023 A Dataset for Audio-Visual Sound Event Detection in Movies
abstract
Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at scale. Movies depict various real-life and fictional scenarios which makes them a rich resource for mining a wide range of audio events. In this work, we present a dataset of audio events called Subtitle-Aligned Movie Sounds (SAM-S). We use publicly available closed-caption transcripts to automatically mine over 110K audio events from 430 movies. We identify three dimensions to categorize audio events: sound, source, quality, and present the steps involved to produce a final taxonomy of 245 sounds. We discuss the choices involved in generating the taxonomy, and also highlight the human-centered nature of sounds in our dataset. We establish a baseline performance for audio-only sound classification of 34.76% mean average precision, and show that incorporating visual information can further improve the performance by 5%. Data and code are made available for research at https://github.com/usc-sail/mica-subtitle-aligned-movie-sounds
Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shri Narayanan
ICASSP2
2023 FedMultimodal: A Benchmark for Multimodal Federated Learning
abstract
Over the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal.
Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan
KDD2
2023 MM-AU: Towards Multimodal Understanding of Advertisement Videos
abstract
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches.
Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan
ACM Multimedia1
2023 SEAR: Semantically-grounded Audio Representations
abstract
Audio supports visual story-telling in movies through the use of different sounds. These sounds are often tied to different visual elements, including foreground entities, the interactions between them as well as background context. Visual captions provide a condensed view of an image, providing a natural language description of entities and the relationships between them. In this work, we utilize visual captions to semantically ground audio representations in a self-supervised setup. We leverage state-of-the-art vision-language models to augment movie datasets with visual captions at scale to the order of 9.6M captions to learn audio representations from over 2500 hours of movie data. We evaluate the utility of the learned representations and show state-of-the art performance on two movie understanding tasks, genre and speaking-style classification, outperforming video based methods and audio baselines. Finally, we show that the learned model can be transferred in a zero-shot manner through application in both movie understanding tasks and general action recognition.
Rajat Hebbar, Digbalay Bose, Shri Narayanan
ACM Multimedia2
2023 MovieCLIP: Visual Scene Recognition in Movies
abstract
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don’t consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shri Narayanan
WACV1
2014 Optimal filter design using an improved artificial bee colony algorithm
Digbalay Bose, Subhodip Biswas, Athanasios V. Vasilakos, Sougata Laha
Inf. Sci.1
2013 Migrating forager population in a multi-population Artificial Bee Colony algorithm with modified perturbation schemes
abstract
Swarm Intelligent algorithms focus on imbibing the collective intelligence of a group of simple agents that can work together as a unit. This research article focus on a recently proposed swarm-based metaheuristic called the Artificial Bee Colony (ABC) algorithm and suggests modifications to the algorithmic framework in order to enhance its performance. The proposed ABC variant shall be referred to as MsABC_Fm (Multi swarm Artificial Bee Colony with Forager migration). MsABC_Fm maintains multiple swarm populations that apply different perturbation strategies and gradually migration of the population from worse performing strategy to the better mode of perturbation is promoted. To evaluate the performance of the algorithm, we conduct comparative study involving 8 algorithms and test the problems on 25 benchmark problems proposed in the Special Session on IEEE Congress on Evolutionary Competition 2005. The superiority of the MsABC_Fm approach is also highlighted statistically.
Subhodip Biswas, Souvik Kundu 0001, Digbalay Bose, Swagatam Das, Ponnuthurai N. Suganthan, Bijaya K. Panigrahi
SIS3