Arun Balaji Buduru

dblp:128/4273 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-6267-8138ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 17 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition
abstract
Acoustic side-channel attacks (ASCA) on keyboards pose a significant security risk, as keystrokes can be inferred from the typing acoustics, revealing sensitive information on laptops. Prior studies on ASCA are limited due to dataset constraints, which include a smaller number of users, keyboards, and environments. This limits the exploration of the potential of this attack vector across different keyboards, users, microphones, and ambient noise conditions.
Bikrant Bikram Pratap Maurya, Nitin Choudhury, Daksh Agarwal, Arun Balaji Buduru
AsiaCCS4
2026 How Well do LLMs Assist Parents in Assessing Child Appropriateness of Videos?
abstract
Children’s entertainment has become increasingly digital, with much of it available on video-sharing platforms. Although traditional media such as movies and TV are manually curated for appropriateness, the sheer quantity of videos being uploaded online makes this approach impractical. Current automated techniques fail to capture the diversity in parental supervision caused by varying parental preferences, culture, and other factors, while also lacking the transparency and explainability necessary to build parental trust. This study seeks to evaluate LLM’s ability to assess the appropriateness of videos for children under the age of 7 in an explainable manner and its overall alignment with parental values. Our study shows that while LLMs are less effective at determining appropriateness themselves, they can provide beneficial descriptions of the videos and effectively aid in the parental decision-making process.
Sabila Nawshin, Ashley Phoebe Ishoel, Arun Balaji Buduru, Apu Kapadia
CHI3
2025 Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
abstract
In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model (FM) representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets.
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Sishir Kalita, Arun Balaji Buduru, Rajesh Sharma 0002, S. R. Mahadeva Prasanna
ICASSP6
2025 PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Jaya Sai Kiran Patibandla, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH6
2025 SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Abu Osama Siddiqui, Priyabrata Mallick, Jaya Sai Kiran Patibandla, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH10
2025 Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH7
2025 Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH7
2025 HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH6
2025 Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH6
2025 Towards Machine Unlearning for Paralinguistic Speech Processing
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Vandana Rajan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH8
2024 Whispers of Trauma: Leveraging Social Media for Assessing Mental Health in Victims of Childhood Sexual Abuse
Orchid Chetia Phukan, Rajesh Sharma 0002, Arun Balaji Buduru
ASONAM (4)3
2024 Understanding Coordinated Communities through the Lens of Protest-Centric Narratives: A Case Study on #CAA Protest
abstract
Social media platforms, particularly Twitter, have emerged as vital media for organizing online protests worldwide. During protests, users on social media share different narratives, often coordinated to share collective opinions and obtain widespread reach. In this paper, we focus on the communities formed during a protest and the collective narratives they share, using the protest on the enactment of the Citizenship Amendment Act (#CAA) by the Indian Government as a case study. Since #CAA protest led to divergent discourse in the country, we first classify the users into opposing stances, i.e., protesters (who opposed the Act) and counter-protesters (who supported it) in an unsupervised manner. Next, we identify the coordinated communities in the opposing stances and examine the collective narratives shared by coordinated communities of opposing stances. We use content-based metrics to identify user coordination, including hashtags, mentions, and retweets. Our results suggest mention as the strongest metric for coordination across the opposing stances. Next, we decipher the collective narratives in the opposing stances using an unsupervised narrative detection framework and found call-to-action, on-ground activity, grievances sharing, questioning, and skepticism narratives in the protest tweets. We analyze the strength of the different coordinated communities using network measures, and perform inauthentic activity analysis on the most coordinated communities on both sides. Our findings also suggest that coordinated communities, which were highly inauthentic, showed the highest clustering coefficient towards a greater extent of coordination.
Kumari Neha 0001, Vibhu Agrawal, Saurav Chhatani, Rajesh Sharma 0002, Arun Balaji Buduru, Ponnurangam Kumaraguru
ICWSM5
2024 FGA: Fourier-Guided Attention Network for Crowd Count Estimation
abstract
Crowd counting is gaining societal relevance, particularly in domains of Urban Planning, Crowd Management, and Public Safety. This paper introduces Fourier-guided attention (FGA), a novel attention mechanism for crowd count estimation designed to address the inefficient full-scale global pattern capture in existing works on convolution-based attention networks. FGA efficiently captures multi-scale information, including full-scale global patterns, by utilizing Fast-Fourier Transformations (FFT) along with spatial attention for global features and convolutions with channel-wise attention for semi-global and local features. The architecture of FGA involves a dual-path approach: (1) a path for processing full-scale global features through FFT, allowing for efficient extraction of information in the frequency domain, and (2) a path for processing remaining feature maps for semi-global and local features using traditional convolutions and channel-wise attention. This dual-path architecture enables FGA to seamlessly integrate frequency and spatial information, enhancing its ability to capture diverse crowd patterns. We apply FGA in the last layers of two popular crowd-counting works, CSRNet and CANNet, to evaluate the module’s performance on benchmark datasets such as ShanghaiTech-A, ShanghaiTech-B, UCF-CC-50, and JHU++ crowd. The experiments demonstrate a notable improvement across all datasets based on Mean-Squared-Error (MSE) and Mean-Absolute-Error (MAE) metrics, showing comparable performance to recent state-of-the-art methods. Additionally, we illustrate the interpretability using qualitative analysis, leveraging Grad-CAM heatmaps, to show the effectiveness of FGA in capturing crowd patterns.
Yashwardhan Chaudhuri, Arun Balaji Buduru, Adel Alshamrani
IJCNN3
2024 ASGIR: audio spectrogram transformer guided classification and information retrieval for birds
Yashwardhan Chaudhuri, Paridhi Mundra, Arnesh Batra, Orchid Chetia Phukan, Arun Balaji Buduru
INTERSPEECH5
2024 The reasonable effectiveness of speaker embeddings for violence detection
Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH3
2024 PERSONA: an application for emotion recognition, gender recognition and age estimation
Devyani Koshal, Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH4
2024 VoxMed: one-step respiratory disease classifier using digital stethoscope sounds
Paridhi Mundra, Manik Sharma, Yashwardhan Chaudhuri, Orchid Chetia Phukan, Arun Balaji Buduru
INTERSPEECH5
2024 ComFeAT: combination of neural and spectral features for improved depression detection
Orchid Chetia Phukan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH5
2024 Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
Orchid Chetia Phukan, Gautam Siddharth Kashyap, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH3
2024 Towards Multilingual Audio-Visual Question Answering
Orchid Chetia Phukan, Priyabrata Mallick, Swarup Ranjan Behera, Aalekhya Satya Narayani, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH5
2024 AVR: synergizing foundation models for audio-visual humor detection
Sarthak Sharma, Orchid Chetia Phukan, Drishti Singh, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH4
2023 Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks
Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma 0002
INTERSPEECH2
2022 The Pursuit of Being Heard: An Unsupervised Approach to Narrative Detection in Online Protest
abstract
Protests and mass mobilization are scarce; however, they may lead to dramatic outcomes when they occur. Social media such as Twitter has become a center point for the organization and development of online protests worldwide. It becomes crucial to decipher various narratives shared during an online protest to understand people's perceptions. In this work, we propose an unsupervised clustering-based framework to understand the narratives present in a given online protest. Through a comparative analysis of tweet clusters in 3 protests around government policy bills, we contribute novel insights about narratives shared during an online protest. Across case studies of government policy-induced online protests in India and the United Kingdom, we found familiar mass mo-bilization narratives across protests. We found reports of on-ground activities and call-to-action for people's participation narrative clusters in all three protests under study. We also found protest-centric narratives in different protests, such as skepticism around the topic. The results from our analysis can be used to understand and compare people's perceptions of future mass mobilizations.
Kumari Neha 0001, Vibhu Agrawal, Arun Balaji Buduru, Ponnurangam Kumaraguru
ASONAM3
2022 Understanding the Impact of Awards on Award Winners and the Community on Reddit
abstract
Non-financial incentives in the form of awards often act as a driver of positive reinforcement and elevation of social status in the offline world. The elevated social status results in people becoming more active, aligning to a change in the communities' expectations. However, the impact in terms of longevity of social influence and community acceptance of leaders of these incentives in the form of awards are not well-understood in the online world. Our work aims to shed light on the impact of these awards on the awardee and the community. We focus on three large subreddits with a snapshot of 219K posts and 5.8 million comments contributed by 88K Reddit users who received 14,146 awards. Our work establishes that the behaviour of awardees change statistically significantly for a short time after getting an award; however, the change is ephemeral since the awardees return to their pre-award behaviour within days. Additionally, via a user survey, we identified a long-lasting impact of awards-we found that the community's stance softened towards awardees.
Avinash Tulasi, Mainack Mondal, Arun Balaji Buduru, Ponnurangam Kumaraguru
ASONAM3
2021 Truth and travesty intertwined: a case study of #SSR counterpublic campaign
abstract
Twitter has emerged as a prominent social media platform for activism and counterpublic narratives. The counterpublics leverage hashtags to build a diverse support network and share content on a global platform that counters the dominant narrative. This paper applies the framework of connective action on the counter-narrative campaign over the cause of death of #SushantSinghRajput. We combine descriptive network, modularity, and hashtag based topical analysis to identify three major mechanisms underlying the campaign: generative role taking, hashtag-based narratives and formation of alignment network towards a common cause. Using the case study of #SushantSinghRajput, we highlight how connective action framework can be used to identify different strategies adopted by counterpublics for the emergence of connective action.
Kumari Neha 0001, Tushar Mohan, Arun Balaji Buduru, Ponnurangam Kumaraguru
ASONAM3
2015 Protecting Critical Cloud Infrastructures with Predictive Capability
abstract
Emerging trends in cyber system security breaches, including those in critical infrastructures involving cloud systems, such as in applications of military, homeland security, finance, utilities and transportation systems, have shown that attackers have abundant resources, including both human and computing power, to launch attacks. The sophistication and resources used in attacks reflect that the attackers may be supported by large organizations and in some cases by foreign governments. Hence, there is an urgent need to develop intelligent cyber defense approaches to better protecting critical cloud infrastructures. In order to have much better protection for critical cloud infrastructures, effective approaches with predictive capability are needed. Much research has been done by applying game theory to generating adversarial models for predictive defense of critical infrastructures. However, these approaches have serious limitations, some of which are due to the assumptions used in these approaches, such as rationality and Nash equilibrium, which may not be valid for current and emerging cloud infrastructures. Another major limitation of these approaches is that they do not capture probabilistic human behaviors accurately, and hence do not incorporate human behaviors. In order to greatly improve the protection of critical cloud infrastructures, it is necessary to predict potential security breaches on critical cloud infrastructures with accurate system-wide causal relationship and probabilistic human behaviors. In this paper, the challenges and our vision on developing such proactive protection approaches are discussed.
Stephen S. Yau, Arun Balaji Buduru, Vinjith Nagaraja
CLOUD2
2015 An Effective Approach to Continuous User Authentication for Touch Screen Smart Devices
abstract
Due to the rapid increase in the use of personal smart devices, more sensitive data is stored and viewed on these smart devices. This trend makes it easier for attackers to access confidential data by physically compromising (including stealing) these smart devices. Currently, most personal smart devices employ one of the one-time user authentication schemes, such as four-to-six digits, fingerprint or pattern-based schemes. These authentication schemes are often not good enough for securing personal smart devices because the attackers can easily extract all the confidential data from the smart device by breaking such schemes, or by keeping the authenticated session open on a physically compromised smart device. In addition, existing re-authentication or continuous authentication techniques for protecting personal smart devices use centralized architecture and require servers at a centralized location to train and update the learning model used for continuous authentication, which impose additional communication overhead. In this paper, an approach is presented to generating and updating the authentication model on the user's smart device with user's gestures, instead of a centralized server. There are two major advantages in this approach. One is that this approach continuously learns and authenticates finger gestures of the user in the background without requiring the user to provide specific gesture inputs. The other major advantage is to have better authentication accuracy by treating uninterrupted user finger gestures over a short time interval as a single gesture for continuous user authentication.
Arun Balaji Buduru, Stephen S. Yau
QRS1