Jagabandhu Mishra

dblp:213/9580 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-1878-6286ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Towards explainable spoofed speech attribution and detection: A probabilistic approach for characterizing speech synthesizer components
abstract
We propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed probabilistic attribute embeddings aim to detect specific speech synthesizer components, represented through high-level attributes and their corresponding values. We use these probabilistic embeddings with four classifier back-ends to address two downstream tasks: spoofing detection and spoofing attack attribution. The former is the well-known bonafide-spoof detection task, whereas the latter seeks to identify the source method (generator) of a spoofed utterance. We additionally use Shapley values, a widely used technique in machine learning, to quantify the relative contribution of each attribute value to the decision-making process in each task. Results on the ASVspoof2019 dataset demonstrate the substantial role of waveform generator, conversion model outputs, and inputs in spoofing detection; and inputs, speaker, and duration modeling in spoofing attack attribution. In the detection task, the probabilistic attribute embeddings achieve 99.7% balanced accuracy and 0.22% equal error rate (EER), closely matching the performance of raw embeddings (99.9% balanced accuracy and 0.22% EER). Similarly, in the attribution task, our embeddings achieve 90.23% balanced accuracy and 2.07% EER, compared to 90.16% and 2.11% with raw embeddings. These results demonstrate that the proposed framework is both inherently explainable by design and capable of achieving performance comparable to raw CM embeddings.
Jagabandhu Mishra, Manasi Chhibber, Hye-Jin Shim, Tomi Kinnunen
Comput. Speech Lang.1
2025 An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech Characterization
abstract
We propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not easy to interpret, the probabilistic attributes are designed to gauge the presence or absence of sub-components that make up a specific spoofing attack. These attributes are then applied to two downstream tasks: spoofing detection and attack attribution. To enforce interpretability also to the back-end, we adopt a decision tree classifier. Our experiments on the ASVspoof2019 dataset with spoof CM embeddings extracted from three models (AASIST, Rawboost-AASIST, SSL-AASIST) suggest that the performance of the attribute embeddings are on par with the original raw spoof CM embeddings for both tasks. The best performance achieved with the proposed approach for spoofing detection and attack attribution, in terms of accuracy, is 99.7% and 99.2%, respectively, compared to 99.7% and 94.7% using the raw CM embeddings. To analyze the relative contribution of each attribute, we estimate their Shapley values. Attributes related to acoustic feature prediction, waveform generation (vocoder), and speaker modeling are found important for spoofing detection; while duration modeling, vocoder, and input type play a role in spoofing attack attribution.
Manasi Chhibber, Jagabandhu Mishra, Hye-Jin Shim, Tomi Kinnunen
ICASSP2
2025 STOPA: A Dataset of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
abstract
A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introduce STOPA, a systematically varied and metadata-rich dataset for deepfake speech source tracing, covering 8 AMs, 6 VMs, and diverse parameter settings across 700k samples from 13 distinct synthesisers. Unlike existing datasets, which often feature limited variation or sparse metadata, STOPA provides a systematically controlled framework covering a broader range of generative factors, such as the choice of the vocoder model, acoustic model, or pretrained weights, ensuring higher attribution reliability. This control improves attribution accuracy, aiding forensic analysis, deepfake detection, and generative model transparency.
Anton Firc, Manasi Chhibber, Jagabandhu Mishra, Vishwanath Pratap Singh, Tomi Kinnunen, Kamil Malinka
INTERSPEECH3
2025 Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
Parismita Gogoi, Vishwanath Pratap Singh, Seema Khadirnaikar, Soma Siddhartha, Sishir Kalita, Jagabandhu Mishra, Md. Sahidullah, Priyankoo Sarmah, S. R. Mahadeva Prasanna
INTERSPEECH6
2025 Optimizing a-DCF for Spoofing-Robust Speaker Verification
abstract
Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. We propose a spoofing-robust ASV system optimized directly for the recently introduced architecture-agnostic detection cost function (a-DCF), which allows targeting a desired trade-off between the contradicting aims of user convenience and robustness to spoofing. We combine a-DCF and binary cross-entropy (BCE) with a novel straightforward threshold optimization technique. Our results with an embedding fusion system on ASVspoof2019 data demonstrate relative improvement of 13% over a system trained using BCE only (from minimum a-DCF of 0.1445 to 0.1254). Using an alternative non-linear score fusion approach provides relative improvement of 43% (from minimum a-DCF of 0.0508 to 0.0289).
Oguzhan Kurnaz, Jagabandhu Mishra, Tomi Kinnunen, Cemal Hanilçi
IEEE Signal Process. Lett.2
2024 Implicit Self-Supervised Language Representation for Spoken Language Diarization
abstract
The use of spoken language diarization (LD) as a preprocessing system might be essential in a code-switched (CS) scenario. Furthermore, implicit frameworks are preferable to explicit ones, as implicit frameworks can be easily adapted to deal with low/zero resource languages. Inspired by speaker diarization literature, three frameworks based on (a) fixed segmentation, (b) change-point-based segmentation, and (c) end-to-end (E2E) are used in this study to perform LD. The initial exploration in the constructed text-to-speech female language diarization (TTSF-LD) dataset shows, that using the x-vector as implicit language representation with appropriate analysis window length achieves, comparable performance to explicit LD. The best implicit LD performance of 6.4% in terms of Jaccard error rate (JER) is achieved by using the E2E framework. However, using the natural Microsoft CS dataset, the performance of the E2E implicit LD degrades to 60.4% JER. The performance degradation is due to the inability of the x-vector representation to capture language-specific traits. To address this shortcoming, a self-supervised implicit language representation framework is used in this study. Compared to the x-vector representation, the self-supervised representation yields a relative improvement of 63.9%, achieving a JER of 21.8% when used in conjunction with the E2E framework.
Jagabandhu Mishra, S. R. Mahadeva Prasanna
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 End to End Spoken Language Diarization with Wav2vec Embeddings
Jagabandhu Mishra, Jayadev N. Patil, Amartya Chowdhury, S. R. Mahadeva Prasanna
INTERSPEECH1
2020 VOP Detection in Variable Speech Rate Condition
Ayush Agarwal, Jagabandhu Mishra, S. R. Mahadeva Prasanna
INTERSPEECH2