Goutam Saha 0001

dblp:s/GoutamKumarSaha · also Goutam Kumar Saha 0001 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0001-6187-1684ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neural style transfer architectures for improving generalization in low-resource spoken language identification
Spandan Dey, Goutam Saha 0001
Eng. Appl. Artif. Intell.2
2024 Improving self-supervised learning model for audio spoofing detection with layer-conditioned embedding fusion
Souvik Sinha, Spandan Dey, Goutam Saha 0001
Comput. Speech Lang.3
2024 Towards Cross-Corpora Generalization for Low-Resource Spoken Language Identification
abstract
Low-resource spoken language identification (LID) systems are prone to poor generalization across unknown domains. In this study, using multiple widely used low-resourced South Asian LID corpora, we conduct an in-depth analysis for understanding the key non-lingual bias factors that create corpora mismatch and degrade LID generalization. To quantify the biases, we extract different data-driven and rule-based summary vectors that capture non-lingual aspects, such as speaker characteristics, spoken context, accents or dialects, recording channels, background noise, and environments. We then conduct a statistical analysis to identify the most crucial non-lingual bias factors and corpora mismatch components that impact LID performance. Following these analyses, we then propose effective bias compensation approaches for the most relevant summary vectors. We generate pseudo-labels using hierarchical clustering over language-domain-gender constrained summary vectors and use them to train adversarial networks with conditioned metric loss. The compensations learn invariance for the corpora mismatches due to the non-lingual biases and help to improve the generalization. With the proposed compensation method, we improve equal error rate up to 5.22% and 8.14% for the same-corpora and cross-corpora evaluations, respectively.
Spandan Dey, Md. Sahidullah, Goutam Saha 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification
Spandan Dey, Premjeet Singh, Goutam Saha 0001
INTERSPEECH3
2023 A novel frequency warping scale for speech emotion recognition
Premjeet Singh, Goutam Saha 0001
INTERSPEECH2
2023 Cross-corpora spoken language identification with domain diversification and generalization
abstract
This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor cross-corpora generalization due to corpora-dependent non-lingual biases. Our contribution to this work is twofold. First, we propose domain diversification, which diversifies the limited training data using different audio data augmentation methods. We then propose the concept of maximally diversity-aware cascaded augmentations and optimize the augmentation fold-factor for effective diversification of the training data. Second, we introduce the idea of domain generalization considering the augmentation methods as pseudo-domains. Towards this, we investigate both domain-invariant and domain-aware approaches. Our LID system is based on the state-of-the-art emphasized channel attention, propagation, and aggregation based time delay neural network (ECAPA-TDNN) architecture. We have conducted extensive experiments with three widely used corpora for Indian LID research. In addition, we conduct a final blind evaluation of our proposed methods on the Indian subset of VoxLingua107 corpus collected in the wild. Our experiments demonstrate that the proposed domain diversification is more promising over commonly used simple augmentation methods. The study also reveals that domain generalization is a more effective solution than domain diversification. We also notice that domain-aware learning performs better for same-corpora LID, whereas domain-invariant learning is more suitable for cross-corpora generalization. Compared to basic ECAPA-TDNN, its proposed domain-invariant extensions improve the cross-corpora EER up to 5.23%. In contrast, the proposed domain-aware extensions also improve performance for same-corpora test scenarios.
Spandan Dey, Md. Sahidullah, Goutam Saha 0001
Comput. Speech Lang.3
2023 Fragment-level classification of ECG arrhythmia using wavelet scattering transform
Sudestna Nahak, Akanksha Pathak, Goutam Saha 0001
Expert Syst. Appl.3
2023 Addressing the semi-open set dialect recognition problem under resource-efficient considerations
Spandan Dey, Goutam Saha 0001
Speech Commun.2
2023 Modulation spectral features for speech emotion recognition using deep neural networks
Premjeet Singh, Md. Sahidullah, Goutam Saha 0001
Speech Commun.3
2022 Localized multiple kernel learning using graph modularity
Lily S. Chamakura, Goutam Saha 0001
Pattern Recognit. Lett.2
2022 An Overview of Indian Spoken Language Recognition from Machine Learning Perspective
abstract
Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction. A front-end LID module helps to improve the performance of many speech-based applications in the multilingual scenario. India is a populous country with diverse cultures and languages. The majority of the Indian population needs to use their respective native languages for verbal interaction with machines. Therefore, the development of efficient Indian spoken language recognition systems is useful for adapting smart technologies in every section of Indian society. The field of Indian LID has started gaining momentum since the early 2000s, mainly due to the development of several standard multilingual speech corpora for the Indian languages. Even though significant research progress has already been made in this field, to the best of our knowledge, there are not many attempts to analytically review them collectively. In this work, we have conducted one of the very first attempts to present a comprehensive review of the Indian spoken language recognition research field. In-depth analysis has been presented to emphasize the unique challenges of low-resource and mutual influences for developing LID systems in the Indian contexts. Several essential aspects of the Indian LID research, such as the detailed description of the available speech corpora, the major research contributions, including the earlier attempts based on statistical modeling to the recent approaches based on different neural network architectures, and the future research trends are discussed. This review work will help assess the state of the present Indian LID research by any active researcher or any research enthusiasts from related fields.
Spandan Dey, Md. Sahidullah, Goutam Saha 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 Ensembled Transfer Learning and Multiple Kernel Learning for Phonocardiogram Based Atherosclerotic Coronary Artery Disease Detection
abstract
Conventional machine learning has paved the way for a simple, affordable, non-invasive approach for Coronary artery disease (CAD) detection using phonocardiogram (PCG). It leaves a scope to explore improvement of performance metrics by fusion of learned representations from deep learning. In this study, we propose a novel, multiple kernel learning (MKL) for their fusion using deep embeddings transferred from pre-trained convolutional neural network (CNN). The proposed MKL, finds optimal kernel combination by maximizing the similarity with ideal kernel and minimizing the redundancy with other basis kernels. Experiments are performed on 960 PCG epochs collected from 40 CAD and 40 normal subjects. The transferred embeddings attain maximum subject-level accuracy of 89.25% with kappa of 0.7850. Later, their fusion with handcrafted features using the proposed MKL gives an accuracy of 91.19% and kappa 0.8238. The study shows the potential of development of high accuracy CAD detection system by using easy to acquire, non-invasive PCG signal.
Akanksha Pathak, Kayapanda Mandana, Goutam Saha 0001
IEEE J. Biomed. Health Informatics3
2021 Robust kernelized graph-based learning
Supratim Manna, Jessy Rimaya Khonglah, Anirban Mukherjee 0001, Goutam Saha 0001
Pattern Recognit.4
2020 Analysis and classification of acoustic scenes with wavelet transform-based mel-scaled features
Shefali Waldekar, Goutam Saha 0001
Multim. Tools Appl.2
2019 An instance voting approach to feature selection
Lily S. Chamakura, Goutam Saha 0001
Inf. Sci.2
2018 Wavelet Transform Based Mel-scaled Features for Acoustic Scene Classification
Shefali Waldekar, Goutam Saha 0001
INTERSPEECH2
2018 Synthetic speech detection using fundamental frequency variation and spectral features
Monisankha Pal, Dipjyoti Paul, Goutam Saha 0001
Comput. Speech Lang.3
2017 Generalization of spoofing countermeasures: A case study with ASVspoof 2015 and BTAS 2016 corpora
abstract
Voice-based biometric systems are highly prone to spoofing attacks. Recently, various countermeasures have been developed for detecting different kinds of attacks such as replay, speech synthesis (SS) and voice conversion (VC). Most of the existing studies are conducted with a specific training set defined by the evaluation protocol. However, for realistic scenarios, selecting appropriate training data is an open challenge for the system administrator. Motivated by this practical concern, this work investigates the generalization capability of spoofing countermeasures in restricted training conditions where speech from a broad attack types are left out in the training database. We demonstrate that different spoofing types have considerably different generalization capabilities. For this study, we analyze the performance using two kinds of features, mel-frequency cepstral coefficients (MFCCs) which are considered as baseline and recently proposed constant Q cepstral coefficients (CQCCs). The experiments are conducted with standard Gaussian mixture model - maximum likelihood (GMM-ML) classifier on two recently released spoofing corpora: ASVspoof 2015 and BTAS 2016 that includes cross-corpora performance analysis. Feature-level analysis suggests that static and dynamic coefficients of spectral features, both are important for detecting spoofing attacks in the real-life condition.
Dipjyoti Paul, Md. Sahidullah, Goutam Saha 0001
ICASSP3
2017 Spectral Mapping Using Prior Re-Estimation of i-Vectors and System Fusion for Voice Conversion
abstract
In this paper, we propose a new voice conversion (VC) method using i-vectors which consider low-dimensional representation of speech utterances. An attempt is made to restrict the i-vector variability in the intermediate computation of total variability (T) matrix by using a novel approach that uses modified-prior distribution of the intermediate i-vectors. This T-modification improves the speaker individuality conversion. For further improvement of conversion score and to keep a better balance between similarity and quality, band-wise spectrogram fusion between conventional joint density Gaussian mixture model (JDGMM) and i-vector based converted spectrograms is employed. The fused spectrogram retains more spectral details and leverages the complementary merits of each subsystem. Experiments in terms of objective and subjective evaluation are conducted extensively on CMU ARCTIC database. The results show that the proposed technique can produce a better trade-off between similarity and quality score than other state-of-the-art baseline VC methods. Furthermore, it works better than JDGMM in limited VC training data. The proposed VC performs moderately better (both objective and subjective) than mixture of factor analyzer based baseline VC. In addition, the proposed VC provides better quality converted speech as compared to maximum likelihood-GMM VC with dynamic feature constraint.
Monisankha Pal, Goutam Saha 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Classification of emotions induced by music videos and correlation with participants' rating
Syed Naser Daimi, Goutam Saha 0001
Expert Syst. Appl.2
2013 A Novel Windowing Technique for Efficient Computation of MFCC for Speaker Recognition
abstract
In this letter, we propose a novel family of windowing technique to compute mel frequency cepstral coefficient (MFCC) for automatic speaker recognition from speech. The proposed method is based on fundamental property of discrete time Fourier transform (DTFT) related to differentiation in frequency domain. Classical windowing scheme such as Hamming window is modified to obtain derivatives of discrete time Fourier transform coefficients. It is mathematically shown that this technique takes into account slope of power spectrum and phase information. Speaker recognition systems based on our proposed family of window functions are shown to attain substantial and consistent performance improvement over baseline single tapered Hamming window as well as recently proposed multitaper windowing technique.
Md. Sahidullah, Goutam Saha 0001
IEEE Signal Process. Lett.2
2012 A system for behavior prediction based on neural signals
Jacob Mathew, Laxmikanta Sahoo, Goutam Saha 0001
Neurocomputing3
2012 Design, analysis and experimental evaluation of block based transformation in MFCC computation for speaker recognition
Md. Sahidullah, Goutam Saha 0001
Speech Commun.2
2010 Detection of cardiac abnormality from PCG signal using LMS based least square SVM classifier
Samit Ari, Koushik Hembram, Goutam Saha 0001
Expert Syst. Appl.3
2010 Feature selection using singular value decomposition and QR factorization with column pivoting for text-independent speaker identification
Sandipan Chakroborty, Goutam Saha 0001
Speech Commun.2
2008 Speech enhancement by joint statistical characterization in the Log Gabor Wavelet domain
Suman Senapati, Sandipan Chakroborty, Goutam Saha 0001
Speech Commun.3
2004 A Comparative Study of Feature Extraction Algorithms on ANN Based Speaker Model for Speaker Recognition Applications
Goutam Saha 0001, Sandipan Chakroborty
ICONIP1
1999 On robust nonlinear modeling of a complex process with large number of inputs using m-QRcp factorization and Cp statistic
abstract
The problem of modeling complex processes with a large number of inputs is addressed. A new method is proposed for the optimization of the models in minimum C(p) statistic sense using QR with a modified scheme of column pivoting (m-QRcp) factorization. Two different classes of multilayer nonlinear modeling problems are explored: 1) in the first class of models, each layer comprises multiple linearly parameterized submodels or cells; the individual cells are optimally modeled using QR factorization, and m-QRcp factorization ensures optimal selection of variables across the layers. 2) The nonhomogeneous feed-forward neural network is chosen as the second class of models, where the network architecture and structure are optimized in terms of best set of hidden links (and nodes) using m-QPcp factorization. In both the cases, the optimization is shown to be direct and conclusive. The proposed is a generic approach to the optimal modeling of complex multilayered architectures, which leads to computationally fast and numerically robust parsimonious designs, free from collinearity problems. The method is largely free from heuristics and is amenable to automated modeling.
Partha Pratim Kanjilal, Goutam Saha 0001, Thomas Jacob Koickal
IEEE Trans. Syst. Man Cybern. Part B2