EDBT 2026 Demo / reviewers in the wild / expert
Swarup Ranjan Behera
dblp:258/8329
· DBLP profile ↗
17ranked-venue papers
4as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion RecognitionabstractIn this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model (FM) representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets. Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Sishir Kalita, Arun Balaji Buduru, Rajesh Sharma 0002, S. R. Mahadeva Prasanna |
ICASSP | 4 |
| 2025 | PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Jaya Sai Kiran Patibandla, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Abu Osama Siddiqui, Priyabrata Mallick, Jaya Sai Kiran Patibandla, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 4 |
| 2025 | Towards Machine Unlearning for Paralinguistic Speech Processing
Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Vandana Rajan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 5 |
| 2024 | FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
Swarup Ranjan Behera, Abhishek Dhiman, Karthik Gowda, Aalekhya Satya Narayani |
INTERSPEECH | 1 |
| 2024 | Towards Multilingual Audio-Visual Question Answering
Orchid Chetia Phukan, Priyabrata Mallick, Swarup Ranjan Behera, Aalekhya Satya Narayani, Arun Balaji Buduru, Rajesh Sharma 0002 |
INTERSPEECH | 3 |
| 2023 | Towards Multi-Lingual Audio Question Answering
Swarup Ranjan Behera, Pailla Balakrishna Reddy, Achyut Mani Tripathi, Megavath Bharadwaj Rathod, Tejesh Karavadi |
INTERSPEECH | 1 |
| 2022 | Defensive Bit Planes: Defense Against Adversarial AttacksabstractDeep learning-based models are vulnerable to adversarial examples crafted with different adversarial attack techniques. Numerous attack methods have been proposed that utilize gradient information of deep model to craft an adversarial examples. Amongst the existing defense mechanisms, adversarial training has gained considerable attention in building robust deep models that remain effective against different adversarial attacks. However, adversarial training demands high computational cost during the development of a robust deep model. In this paper, we present a simple yet effective defense mechanism against adversarial attacks. The proposed defense mechanism uses the concept of bit plane slicing for de-noising of an input image. The efficacy of the proposed defense technique has been evaluated on two benchmark image datasets, viz. MNIST and Fashion-MNIST datasets. The experiments and results show that the proposed defence technique yields comparable and competitive performance to state-of-the-art defense techniques against adversarial attacks. Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul |
IJCNN | 2 |
| 2022 | Adv-IFD: Adversarial Attack Datasets for An Intelligent Fault DiagnosisabstractDeep learning techniques have been widely applied for performing intelligent fault diagnosis (IFD) for applications such as bearing fault diagnosis, wind turbines and drilling operations. Notwithstanding the huge success of deep learning-based models for intelligent fault diagnosis, their vulnerability against adversarial attacks have been neglected to a great extent. The adversarial attacks aim to craft an adversarial sample by addition of imperceptible perturbation to its clean data sample. In this paper, we investigate the performance of different deep models against four state-of-the-art adversarial attacks. A total of four deep models have been tested against untargeted white-box adversarial attacks. Moreover, analysis of transferability of adversarial examples across different deep models is also inspected. Experiments and results reveal that deep models for IFD are highly susceptible to adversarial examples crafted with four state-of-the-art adversarial attacks. The proposed work presents an extensive insight on adversarial samples of machinery vibration signals of the CWRU dataset. Additionally, we are also releasing a first ever adversarial attack dataset for IFD, i.e. Adv-IFD. The code used in this work and adversarial attack datasets are available at: https://github.com/achyutmani/ADV-IFD Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul |
IJCNN | 2 |
| 2022 | Reverse Adversarial Attack To Enhance Environmental Sound ClassificationabstractClassification of environmental sounds becomes an essential unit for several applications such as robot audition, audio-visual and ambient intelligent systems. Recent years have witnessed remarkable success of adversarial attacks for misguiding deep models with a high fooling rate. Adversarial examples crafted with adversarial techniques pose serious security threats for deployment of deep models to practical scenarios. In this paper, we show that adversarial signals crafted by performing search in the opposite direction of the search direction of the perturbations could be utilized while training of deep classifiers of environmental sound classification (ESC). The proposed reverse adversarial attack (RAA) technique can be used as data augmentation tool to enhance the accuracy of deep model that performs ESC. The deep network trained with the crafted adversarial signal shows high robustness against adversarial attacks. The efficacy of the proposed method is evaluated on two benchmark ESC datasets viz. ESC-10 and DCASE-2019 Task-1A datasets. The experiments and results show that the model trained with the proposed technique shows comparable and competitive performance to state-of-the-art (SOTA) deep ESC models. Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul |
IJCNN | 2 |
| 2022 | Investigation of Performance of Visual Attention Mechanisms for Environmental Sound Classification: A Comparative StudyabstractClassification of environmental sound is an equally challenging task to classification of human speech or music. In recent years, attention mechanism has gained considerable attention in learning representative and prototypical features to resolve various research challenges from domains such as computer vision, text mining and sound classification. Advancements in the computer vision can be used for improving the performance of deep models designed to perform environmental sound classification as well. In this paper, we investigated the performance of deep network when twelve SOTA visual attention mechanisms are incorporated in the training of the deep network. The performance of the deep model is evaluated on two benchmark environmental sound classification datasets, viz. ESC-10 and DCASE-2019 task-1(A) datasets. Achyut Mani Tripathi, Swarup Ranjan Behera, Konark Paul |
IJCNN | 2 |
| 2020 | Mining Temporal Changes in Strengths and Weaknesses of Cricket Players Using Tensor Decomposition
Swarup Ranjan Behera, Vijaya Saradhi Vedula |
ESANN | 1 |
| 2019 | Mining Strengths and Weaknesses of Cricket Players Using Short Text CommentaryabstractKnowledge of strengths and weaknesses of players is the key for team selection and strategy planning in any team sport such as Cricket. Computationally, this problem is mostly unexplored. Existing methods focus only on aggregate and macroscopic statistics that ignore many details. The central idea of our paper is to mine strength and weakness rules using short text commentary data. This dataset is compact, semi-structured, accurate, and yet ignored by the machine learning community until now. We collect fine-grained information about each player from the short text commentary dataset and represent it using domain-specific features identified by us. We employ a dimensionality reduction method specific to discrete random variable case, namely correspondence analysis and construct semantic relation between bowler and batsman. This relation is plotted using biplots. Human readable strength and weakness rules are extracted from the biplots. We have performed experiments using a large dataset that describes over one million deliveries. We validate our extracted rules using both intrinsic and extrinsic validation. Swarup Ranjan Behera, Parag Agrawal, Amit Awekar, Vijaya Saradhi Vedula |
ICMLA | 1 |