Swapnil Bhosale

dblp:246/3229 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0001-8299-7569ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unsupervised Audio-Visual Segmentation with Modality Alignment
abstract
Audio-Visual Segmentation (AVS) aims to identify, at the pixel level, the object in a visual scene that produces a given sound. Current AVS methods rely on costly fine-grained annotations of mask-audio pairs, making them impractical for scalability. To address this, we propose the Modality Correspondence Alignment (MoCA) framework, which seamlessly integrates off-the-shelf foundation models like DINO, SAM, and ImageBind. Our approach leverages existing knowledge within these models and optimizes their joint usage for multimodal associations. Our approach relies on estimating positive and negative image pairs in the feature space. For pixel-level association, we introduce an audio-visual adapter and a novel {pixel matching aggregation} strategy within the image-level contrastive learning framework. This allows for a flexible connection between object appearance and audio signal at the pixel level, with tolerance to imaging variations such as translation and rotation. Extensive experiments on the AVSBench (single and multi-object splits) and AVSS datasets demonstrate that MoCA outperforms unsupervised baseline approaches and some supervised counterparts, particularly in complex scenarios with multiple auditory objects. In terms of mIoU, MoCA achieves a substantial improvement over baselines in both the AVSBench (S4: +17.24%, MS3: +67.64%) and AVSS (+19.23%) audio-visual segmentation challenges.
Swapnil Bhosale, Haosen Yang 0003, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
AAAI1
2024 DiffSED: Sound Event Detection with Denoising Diffusion
abstract
Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the split-and-classify (i.e., frame-level) strategy or the more principled event-level modeling approach, all existing methods consider the SED problem from the discriminative learning perspective. In this work, we reformulate the SED problem by taking a generative learning perspective. Specifically, we aim to generate sound temporal boundaries from noisy proposals in a denoising diffusion process, conditioned on a target audio sample. During training, our model learns to reverse the noising process by converting noisy latent queries to the ground-truth versions in the elegant Transformer decoder framework. Doing so enables the model generate accurate event boundaries from even noisy queries during inference. Extensive experiments on the Urban-SED and EPIC-Sounds datasets demonstrate that our model significantly outperforms existing alternatives, with 40+% faster convergence in training. Code: https://github.com/Surrey-UPLab/DiffSED
Swapnil Bhosale, Sauradip Nag, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
AAAI1
2024 Product Retrieval and Ranking for Alphanumeric Queries
abstract
This talk addresses the challenge of improving user experience on e-commerce platforms by enhancing product ranking relevant to user's search queries. Queries such as S2716DG consist of alphanumeric characters where a letter or number can signify important detail for the product/model. Speaker describes recent research where we curate samples from existing datasets at eBay, manually annotated with buyer-centric relevance scores, and centrality scores which reflect how well the product title matches the user's intent. We introduce a User-intent Centrality Optimization (UCO) approach for existing models, which optimizes for the user intent in semantic product search. To that end, we propose a dual-loss based optimization to handle hard negatives, i.e., product titles that are semantically relevant but do not reflect the user's intent. Our contributions include curating a challenging evaluation set and implementing UCO, resulting in significant improvements in product ranking efficiency, observed for different evaluation metrics. Our work aims to ensure that the most buyer-centric titles for a query are ranked higher, thereby, enhancing the user experience on e-commerce platforms.
Hadeel Saadany, Swapnil Bhosale, Samarth Agrawal, Constantin Orasan, Diptesh Kanojia
CIKM2
2024 AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
abstract
Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in addition to low efficiency originating from heavy NeRF rendering, these methods all have a limited ability of characterizing the entire scene environment such as room geometry, material properties, and the spatial relation between the listener and sound source. To address these issues, we propose a novel Audio-Visual Gaussian Splatting (AV-GS) model. To obtain a material-aware and geometry-aware condition for audio synthesis, we learn an explicit point-based scene representation with audio-guidance parameters on locally initialized Gaussian points, taking into account the space relation from the listener and sound source. To make the visual scene model audio adaptive, we propose a point densification and pruning strategy to optimally distribute the Gaussian points, with the per-point contribution in sound propagation (e.g., more points needed for texture-less wall surfaces as they affect sound path diversion). Extensive experiments validate the superiority of our AV-GS over existing alternatives on the real-world RWAS and simulation-based SoundSpaces datasets. Project page: \url{https://surrey-uplab.github.io/research/avgs/}
Swapnil Bhosale, Haosen Yang 0003, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
NeurIPS1
2023 A Novel Metric For Evaluating Audio Caption Similarity
abstract
Automatic Audio Captioning (AAC) refers to the task of describing an audio sample in a natural language (NL) text. Unlike NL text generation tasks, which rely on lexical semantic metrics like BLEU for evaluation, the AAC evaluation metric requires acoustic semantics to map NL text corresponding to similar sounds in addition to lexical semantics. In this paper, we propose a novel metric based on Text-to-Audio Grounding (TAG), to incorporate acoustic semantics. Experiments demonstrate our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature for AAC.
Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu
ICASSP1
2021 Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event Detection
abstract
Lung Auscultation is a non-invasive process of distinguishing normal respiratory sounds from abnormal ones by analyzing the airflow along the respiratory tract. With developments in the Deep Learning (DL) techniques and wider access to anonymized medical data, automatic detection of specific sounds such as crackles and wheezes have been gaining popularity. In this paper, we propose to use two sets of diversified acoustic biomarkers extracted using Discrete Wavelet Transform (DWT) and deep encoded features from the intermediate layer of a pre-trained Audio Event Detection (AED) model trained using sounds from daily activities. First set of biomarkers highlight the time frequency localization characteristics obtained from DWT coefficients. However, the second set of deep encoded biomarkers captures a generalized reliable representation, and thus indemnifies the scarcity of training samples and the class imbalance in dataset. The model trained using these features achieves a 15.05% increase in terms of the specificity over the baseline model that uses spectrogram features. Moreover, ensemble of DWT features and deep encoded feature based models show absolute improvements of 8.32%, 6.66% and 7.40% in terms of sensitivity, specificity and ICBHI-score, respectively, and clearly outperforms the state-of-the-art with a significant margin.
Upasana Tiwari, Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu
ICASSP2
2021 Contrastive Learning of Cough Descriptors for Automatic COVID-19 Preliminary Diagnosis
abstract
Cough sounds as a descriptor have been used for detecting various respiratory ailments based on its intensity, duration of intermediate phase between two cough sounds, repetitions, dryness etc.However, COVID-19 diagnosis using only cough sounds is challenging because of cough being a common symptom among many non COVID-19 health diseases and inherent data imbalance within the available datasets.As one of the approach in this direction, we explore the robustness of multi-domain representation by performing the early fusion over a wide set of temporal, spectral and tempo-spectral handcrafted features, followed by training a Support Vector Machine (SVM) classifier.In our second approach, using a contrastive loss function we learn a latent space from Mel Filter Cepstral Coefficients (MFCCs) where representations belonging to samples having similar cough characteristics are closer.This helps learn representations for the highly varied COVIDnegative class (healthy and symptomatic COVID-negative), by learning multiple smaller clusters.Using only the DiCOVA data, multi-domain features yields an absolute improvement of 0.74% and 1.07%, whereas our second approach shows an improvement of 2.09% and 3.98%, over the blind test and validation set, respectively, when compared with challenge baseline.
Swapnil Bhosale, Upasana Tiwari, Rupayan Chakraborty, Sunil Kumar Kopparapu
Interspeech1
2021 Automatic speaker independent dysarthric speech intelligibility assessment system
Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu
Comput. Speech Lang.2
2020 Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion Recognition
abstract
An End-to-End model with convolutional layers and multi-head self attention mechanism is proposed for Speech Emotion Recognition (SER) task. As inputs, we propose to use both the deep encoded linguistic features that carry the language related context of emotion and the audio spectrogram that are representatives of acoustic cues. To facilitate the deep linguistic feature representation, we use outputs from the intermediate layers of a pre-trained Automatic Speech Recognition (ASR) model, where the layer is selected empirically. The influence of both acoustic and linguistic features, both separately and in combination, for emotion recognition in different scenarios (scripted and spontaneous recording of emotional speech samples) have been studied. Extensive experiments on the standard IEMOCAP database are conducted to investigate the efficacy of our proposed approach. To address the class imbalance, we carried out down sampling and ensembling, which further improved the SER accuracy. Overall, we observe that the acoustic features perform best for improvised recordings which is due to the spontaneity in speech with less linguistic correlation. But the linguistic features are found to be effective for the scripted as well as for the combined (scripted and improvised recordings together) scenario that reflects more linguistic information in spoken utterances.
Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu
ICASSP1
2020 Improved Speaker Independent Dysarthria Intelligibility Classification Using Deepspeech Posteriors
abstract
Individuals with dysarthria are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Automatic intelligibility assessment of dysarthric patients allows clinicians diagnose the impact of therapy and medication and also to plan future course of action. Earlier works have concentrated on building speaker dependent machine learning systems for intelligibility assessment, due to limited availability of data. However, a speaker independent assessment system is of greater use by clinicians. Motivated by this observation, we propose a speaker independent intelligibility assessment system which relies on a novel set of features obtained by processing the output of DeepSpeech, an end to end Speech-to-Text engine. All experiments have been performed on the Universal Access Speech database. An accuracy of 53.9% was obtained using Support Vector Machine based four-class classification system for the speaker independent scenario while the accuracy obtained for the speaker dependent scenario is 97.4%.
Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu
ICASSP2
2020 A Novel Approach for Intelligibility Assessment in Dysarthric Subjects
abstract
Dysarthria is a motor speech impairment caused by muscle weakness. Individuals, with this condition, are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Systems that can assess intelligibility of dysarthric speech can help clinicians diagnose the impact of therapy and medication. In the paper, we propose a usable novel method to assess intelligibility of dysarthric speakers. The approach is based on the observation that the performance of a speech recognition engine deteriorates with increase in severity of the disorder. The mismatch between the original word and the recognized string is exploited to compute the dysarthria intelligibility score. Experiments on UA speech corpus show that the computed intelligibility score exhibits a significant correlation with perceptually assessed intelligibility scores. We further show that a small set of words spoken by the dysarthric subject is sufficient to assess the speech intelligibility reliably.
Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu
ICASSP2
2019 End-to-End Spoken Language Understanding: Bootstrapping in Low Resource Scenarios
Swapnil Bhosale, Imran A. Sheikh, Sri Harsha Dumpala, Sunil Kumar Kopparapu
INTERSPEECH1