EDBT 2026 Demo / reviewers in the wild / expert
Hemant A. Patil
dblp:18/830
· DBLP profile ↗
74ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-4068-2005ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 53 · 8 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature-Level Fusion of Source, System, and Fractal Features for Classification of Infant CriesabstractInfant cry analysis offers a non-invasive way to assess early physiological and developmental conditions. In this work, we evaluate multiple feature representations for infant cry classification, including spectral features such as (MFCC, GFCC, BFCC, CFCC) and self-supervised embeddings (HuBERT, Wav2Vec 2.0, XLS-R, WavLM, TERA), and a Cry Specific Feature Fusion (CSFF) combining acoustic–prosodic and nonlinear fractal descriptors. Experiments are conducted on the Baby Chillanto, Baby Chillanto 2.0, and DAIICT infant cry corpora, covering six pathologies. We further analyze class separability using violin plots of the Convex Hull Area. The results provide a comparative view of how different handcrafted, fused, and pretrained representations perform under clean and noisy conditions. Across all noise conditions, XLS-R 300M achieved the best overall performance with an average accuracy of 87.18%, outperforming the lowest-scoring model (GFCC, 63.41%) by 23.77%. Among handcrafted features, the CSFF fusion features performed strongest with an average of 86.09%, exceeding MFCC by 12.44%, GFCC by 22.68%, and CFCC by 14.82%, highlighting the robustness of the fused representation across Babble, Street, and Pink noise. Hiya Chaudhari, Satyam Rana, Hemant A. Patil |
ICPR (9) | 3 |
| 2026 | WINning Against Audio DeepfakesabstractIn this work, we propose the Wavelet Integrated Network (WIN), a novel neural architecture and employ it for Audio Deepfake Detection (ADD). Unlike standard transformer-based models that mix informationacrosstokens via global self-attention, WIN replaces these attention blocks withintra-tokenwavelet atoms to enhance intra-token understanding of the network. These atoms perform localized time-frequency analysis inside each token, capturing spectral patterns that emphasizelocalstructure rather than the global context. We further propose Multi-Head Wavelet (MHW), and demonstrate that the proposed network has a lower number of parameters as compared to the existing baseline systems. In addition, we conduct ablation studies on variousrealandanalyticwavelets to further optimize the network. We compared our work with baseline, and demonstrate$\approx$2 % decrease in pooled EER over 9 competitive datasets for ADD research. All the codes are publicly available athttps://github.com/Arth-Shah/WIN. Arth J. Shah, Aniket Pandey, Hemant A. Patil |
IEEE Signal Process. Lett. | 3 |
| 2024 | FCHiFi-GAN: Aggrandizing Fast Convergence with Batchwise Normalization
Ravindrakumar M. Purohit, Arushi Srivastava, Hemant A. Patil |
ICPR (31) | 3 |
| 2024 | Linear Frequency Residual Cepstral Features for Dysarthria Severity Classification
Aditya Pusuluri, Hemant A. Patil |
ICPR (20) | 2 |
| 2024 | Infant Cry Classification Using Modified Group Delay Cepstral Coefficients
Arth J. Shah, Hiya Chaudhari, Hemant A. Patil |
ICPR (14) | 3 |
| 2024 | Multi-Block U-Net for Wind Noise Reduction in Hearing Aids
Arth J. Shah, Manish Suthar, Hemant A. Patil |
ICPR (27) | 3 |
| 2024 | Morse wavelet transform-based features for voice liveness detection
Priyanka Gupta 0002, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2024 | Modeling musical expectancy via reinforcement learning and directed graphs
Kirtana Phatnani, Hemant A. Patil |
Multim. Tools Appl. | 2 |
| 2024 | CQT-Based Cepstral Features for Classification of Normal vs. Pathological Infant CryabstractInfant cry classification is an important area of research that involves analyzing cry to detect and classify between normalvs. pathological cries. However, signal processing based state-of-the-art feature sets, such as Short-Time Fourier Transform (STFT) representations and Mel Frequency Cepstral Coefficients (MFCC), have been earlier reported for this task.Quasi-periodicsampling of the vocal tract spectrum by high pitch source harmonics results in poor spectral resolution in the STFT and hence, these feature sets fail to produce a satisfactory classification performance. Contrary to the linearly-spaced frequency bins, this study proposes to use geometrically-spaced frequency bins employed in the CQT-based features, namely, Constant Q Cepstral Coefficients (CQCC) to systematically emphasize the required fundamental frequency (F0) and its harmonics (kF0, k ∊ Z) for infant cry classification. For a comprehensive evaluation of the proposed feature set, two datasets have been considered in this work, namely, Baby Chilanto and In-House DA-IICT datasets. The performance of the proposed CQCC feature set is compared against state-of-the-art MFCC, Linear Frequency Cepstral Coefficients (LFCC), and Cepstral feature sets. Experiments were performed using10-fold cross-validation on two traditional classifiers, namely, Gaussian Mixture Model (GMM) and Support Vector Machine (SVM). Our study finds that better results were obtained using CQCC-GMM architecture with classification accuracies of 99.8% and 98.24% on the Baby Chilanto and In-House DA-IICT datasets, respectively. Further, this work also illustrates the effectiveness of the form-invariance property of the CQT over the traditional narrowband STFTbased spectrogram. Furthermore, this study also presents the effect of parameter tuning and parameter dimension of the feature vector. Furthermore, this study presents the first-ever cross-database and combined dataset scenarios with an overall improvement of 1.59% on the proposed CQCC feature set. Additionally, the robustness of CQCC is evaluated under signal degradation conditions with additive babble noise having various Signal-to-Noise Ratio (SNR) levels on both datasets. Next, the performance of the proposed CQCC was compared with the other feature sets using statistical measures, such asF1-score, J-statistics, violin plots, and analysis of latency period for the deployment of the practical system. Finally, this study compares the best obtained results of CQCC with the existing studies on the Baby Chilanto dataset. Hemant A. Patil, Aastha Kachhi, Ankur T. Patil |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Whisper Encoder features for Infant Cry Classification
Monil Charola, Aastha Kachhi, Hemant A. Patil |
INTERSPEECH | 3 |
| 2023 | Whisper Features for Dysarthric Severity-Level Classification
Siddharth Rathod, Monil Charola, Akshat Vora, Yash Jogi, Hemant A. Patil |
INTERSPEECH | 5 |
| 2023 | Replay spoof detection using energy separation based instantaneous frequency estimation from quadrature and in-phase components
Priyanka Gupta 0002, Piyushkumar K. Chodingala, Hemant A. Patil |
Comput. Speech Lang. | 3 |
| 2023 | On significance of constant-Q transform for pop noise detection
Kuldeep Khoria, Ankur T. Patil, Hemant A. Patil |
Comput. Speech Lang. | 3 |
| 2023 | Multiple voice disorders in the same individual: Investigating handcrafted features, multi-label classification algorithms, and base-learners
Sylvio Barbon Junior, Rodrigo Capobianco Guido, Gabriel Aguiar, Everton Jose Santana, Mario Lemes Proença Jr., Hemant A. Patil |
Speech Commun. | 6 |
| 2022 | Constant Q Cepstral coefficients for classification of normal vs. Pathological infant cryabstractClassification of normal vs. pathological infant cry is an interesting and technologically challenging research problem due to quasi-periodic sampling of vocal tract spectrum by high pitch-source harmonics resulting in extremely poor spectral resolution for commonly used spectral features, such as Mel Frequency Cepstral Coefficients (MFCC). To that effect, in this paper, we propose a new approach of feature extraction based on Constant Q Transform (CQT) that is known to have variable spectro-temporal resolution w.r.t Heisenberg’s un-certainty principle in signal processing framework. Further, CQT is also known to preserve form-invariance property (than its Short-Time Fourier Transform (STFT) counterpart)-a desirable attribute of feature descriptors to be invariant w.r.t shape, shift, rotation, and scaling. CQT- based features are then transformed to the cepstral-domain to derive Constant Q Cepstral Coefficients (CQCC), which are then fed to statistical and discriminative classifiers, namely, Gaussian Mixture Model (GMM) and Support Vector Machine (SVM) respectively. CQCC-GMM and CQCC-SVM systems gave relatively better results than MFCC for various experimental evaluation factors for infant cry classification task on widely used and statistically meaningful Baby Chilanto Database. Relatively best performance, in particular, 99.82% accuracy (0.44% EER), is observed for CQCC-GMM system. Hemant A. Patil, Ankur T. Patil, Aastha Kachhi |
ICASSP | 1 |
| 2022 | Improving the potential of Enhanced Teager Energy Cepstral Coefficients (ETECC) for replay attack detection
Ankur T. Patil, Rajul Acharya, Hemant A. Patil, Rodrigo Capobianco Guido |
Comput. Speech Lang. | 3 |
| 2022 | Effectiveness of energy separation-based instantaneous frequency estimation for cochlear cepstral features for synthetic and voice-converted spoofed speech detection
Ankur T. Patil, Hemant A. Patil, Kuldeep Khoria |
Comput. Speech Lang. | 2 |
| 2022 | Voice privacy using CycleGAN and time-scale modification
Gauri P. Prajapati, Dipesh K. Singh, Preet P. Amin, Hemant A. Patil |
Comput. Speech Lang. | 4 |
| 2022 | Music footprint recognition via sentiment, identity, and setting identification
Kirtana Phatnani, Hemant A. Patil |
Multim. Tools Appl. | 2 |
| 2021 | Cross-Teager Energy Cepstral Coefficients for Replay Spoof Detection on Voice AssistantsabstractVoice assistants (VAs) are highly vulnerable to replay attacks, where the impostor plays pre-recorded voice samples to gain an unauthorized access to personalised devices. To that effect, we present an optimal microphone-channel selection scheme using Cross-Teager Energy Operator (CTEO) for spoofed speech detection (SSD) task. Here, a channel refers to the speech signal obtained from a single microphone among the microphone array. The key idea of this work is optimal channel selection based on maximum cross-energies from a multichannel input, which is suitable for SSD task. This newly proposed feature set is named as Cross-Teager Energy Cepstal Coefficients (CTECCmax). The reason be hind maximizing the cross-energies is to identify the distortions in replay speech signal which is added due to intermediate devices. This key idea is also cross-validated by selecting the least estimated cross-energies as feature set CTECCmin. The noticeable improvement in the performance is observed for CTECCmaxover CTECCminfor two classifiers, namely, Gaussian Mixture Model (GMM) and Light Convolutional Neural Network (LCNN). Rajul Acharya, Harsh Kotta, Ankur T. Patil, Hemant A. Patil |
ICASSP | 4 |
| 2021 | Voice Privacy Through x-Vector and CycleGAN-Based Anonymization
Gauri P. Prajapati, Dipesh K. Singh, Preet P. Amin, Hemant A. Patil |
Interspeech | 4 |
| 2021 | Detection of replay spoof speech using teager energy feature cues
Madhu R. Kamble, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2021 | Residual Neural Network precisely quantifies dysarthria severity-level based on short-duration speech segments
Siddhant Gupta, Ankur T. Patil, Mirali Purohit, Mihir Parmar, Maitreya Patel, Hemant A. Patil, Rodrigo Capobianco Guido |
Neural Networks | 6 |
| 2021 | Non-intrusive quality assessment of noise-suppressed speech using unsupervised deep features
Meet H. Soni, Hemant A. Patil |
Speech Commun. | 2 |
| 2020 | Mspec-Net : Multi-Domain Speech Conversion NetworkabstractIn this paper, we present a multi-domain speech conversion technique by proposing a Multi-domain Speech Conversion Network (MSpeC-Net) architecture for solving the less-explored area of Non-Audible Murmur-to-SPeeCH (NAM2-SPCH) conversion. The murmur produced by the speaker and captured by the NAM microphone undergoes speech quality degradation. Hence, NAM2SPCH conversion becomes a necessary and challenging task for improving the intelligibility of NAM signal. MSpeC-Net contains three domain-specific autoencoders. The multiple encoder-decoders are aligned using latent consistency loss in such a way that the desired conversion is achieved by using the source encoder and target decoder only. We have performed zero-pair NAM2SPCH conversion using the interaction between source encoder and the target decoder. We evaluated our proposed method using both objective and subjective evaluations. With a Mean Opinion Score of 3.26 and 3.12 on an average in a direct NAM2SPCH, and an indirect NAM2SPCH (i.e., NAM-to-whisper-to-speech) conversion, respectively. MSpeC-Net achieves the perceptually significant improvement for NAM2SPCH conversion system. Harshit Malaviya, Jui Shah, Maitreya Patel, Jalansh Munshi, Hemant A. Patil |
ICASSP | 5 |
| 2020 | Amplitude and Frequency Modulation-based features for detection of replay Spoof Speech
Madhu R. Kamble, Hemlata Tak, Hemant A. Patil |
Speech Commun. | 3 |
| 2019 | Novel Enhanced Teager Energy Based Cepstral Coefficients for Replay Spoof DetectionabstractReplay attack on voice biometric, refers to the fraudulent attempt made by an imposter to spoof another person's identity by replaying the pre-recorded voice samples in front of an Automatic Speaker Verification (ASV) system. In an attempt to develop countermeasures against replay attack, this paper proposes to use a new feature set, namely, Enhanced Teager Energy Cepstral Coefficients (ETECC) using the recently introduced concept of signal mass. Results obtained on ASVspoof 2017 version 2.0 dataset suggest that the proposed feature set performs better than the original Teager Energy Cepstral Coefficients (TECC) feature set because the Enhanced Teager Energy Operator (ETEO) gives a better estimate of signal's energy as compared to the Teager Energy Operator (TEO). We obtained 53.3% and 51.35% reduction in EER on development and evaluation dataset, respectively, with respect to the baseline system. Rajul Acharya, Hemant A. Patil, Harsh Kotta |
ASRU | 2 |
| 2019 | Analysis of Reverberation via Teager Energy Features for Replay Spoof Speech DetectionabstractThe Automatic Speaker Verification (ASV) systems are vulnerable to spoofing attacks. Detecting replay attack is the challenging Spoof Speech Detection (SSD) task, as several factors are involved during replay mechanism. Hence, it is important to analyze these factors for effective SSD task. This paper introduces the analysis of the replay speech focusing only on the effect of reverberation on the replay speech. The reverberation introduces delay and change in amplitude producing close copies of natural signal that makes natural components inseparable from the replay components and hence, fails to classify the replay speech signal. To that effect, we propose use of Teager Energy Operator (TEO) to compute running estimate of subband energies for replay vs. natural signal. These subband energies are mapped to cepstraldomain to get proposed Teager Energy Cepstral Coefficients (TECC) for replay SSD task. With the TECC feature set, we analyzed the individual performance for all the Relay Configurations (RC) with Gaussian Mixture Model (GMM) as classifier. The experimental results gave lower Equal Error Rate (EER) of 11.73 % with TECC features and further reduced to 10.30 % with score-level fusion of LFCC and TECC features on evaluation dataset of ASVspoof 2017 challenge version 2.0 database. Madhu R. Kamble, Hemant A. Patil |
ICASSP | 2 |
| 2019 | Novel Metric Learning for Non-parallel Voice ConversionabstractObtaining aligned spectral pairs in case of non-parallel data for stand-alone Voice Conversion (VC) technique is a challenging research problem. Unsupervised alignment algorithm, namely, an Iterative combination of a Nearest Neighbor search step and a Conversion step Alignment (INCA) iteratively tries to align the spectral features by minimizing the Euclidean distance metric between the intermediate converted and the target spectral feature vectors. However, the Euclidean distance may not correlate well with the perceptual distance between the two (sound or visual) patterns in a given feature space. In this paper, we propose to learn distance metric using Large Margin Nearest Neighbor (LMNN) technique that gives a minimum distance for the same phoneme uttered by the different speakers and more distance for the different set of phonemes. This learned metric is then used for finding the NN pairs in the INCA. Furthermore, we propose to use this learned metric only for the first iteration in the INCA, since the intermediate converted features (which are not the actual acoustic features) may not behave well w.r.t. the learned metric. We obtained on an average 7.93 % relative improvement in Phonetic Accuracy (PA). This is reflected positively in subjective and objective evaluations. Nirmesh J. Shah, Hemant A. Patil |
ICASSP | 2 |
| 2019 | Energy Separation-Based Instantaneous Frequency Estimation for Cochlear Cepstral Feature for Replay Spoof Detection
Ankur T. Patil, Rajul Acharya, Pulikonda Krishna Aditya Sai, Hemant A. Patil |
INTERSPEECH | 4 |
| 2019 | Phone Aware Nearest Neighbor Technique Using Spectral Transition Measure for Non-Parallel Voice Conversion
Nirmesh J. Shah, Hemant A. Patil |
INTERSPEECH | 2 |
| 2019 | Whether to Pretrain DNN or not?: An Empirical Analysis for Voice Conversion
Nirmesh J. Shah, Hardik B. Sailor, Hemant A. Patil |
INTERSPEECH | 3 |
| 2019 | Vocal Tract Length Normalization using a Gaussian mixture model framework for query-by-example spoken term detection
Maulik C. Madhavi, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2019 | A novel approach to remove outliers for parallel voice conversion
Nirmesh J. Shah, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2018 | Time-Frequency Masking-Based Speech Enhancement Using Generative Adversarial NetworkabstractThe success of time-frequency (T-F) mask-based approaches is dependent on the accuracy of predicted mask given the noisy spectral features. The state-of-the-art methods in T- F masking-based enhancement employ Deep Neural Network (DNN) to predict mask. Recently, Generative Adversarial Networks (GAN) are gaining popularity instead of maximum likelihood (ML)-based optimization of deep learning architectures. In this paper, we propose to exploit GAN in T-F masking-based enhancement framework. We present the viable strategy to use GAN in such application by modifying the existing approach. To achieve this, we use a method that learns the mask implicitly while predicting the clean T-F representation. Moreover, we show the failure of vanilla GAN in predicting the accurate mask and propose a regularized objective function with the use of Mean Square Error (MSE) between predicted and target spectrum to overcome it. The objective evaluation of the proposed method shows the improvement in the accurate mask prediction, as against the state-of-the-art ML-based optimization techniques. The proposed system significantly improves over a recent GAN-based speech enhancement system in improving speech quality, while maintaining a better trade-off between less speech distortion and more effective removal of background interferences present in the noisy mixture. Meet H. Soni, Neil Shah, Hemant A. Patil |
ICASSP | 3 |
| 2018 | Novel Variable Length Energy Separation Algorithm Using Instantaneous Amplitude Features for Replay Detection
Madhu R. Kamble, Hemant A. Patil |
INTERSPEECH | 2 |
| 2018 | Effectiveness of Speech Demodulation-Based Features for Replay Detection
Madhu R. Kamble, Hemlata Tak, Hemant A. Patil |
INTERSPEECH | 3 |
| 2018 | DA-IICT/IIITV System for Low Resource Speech Recognition Challenge 2018
Hardik B. Sailor, Maddala Venkata Siva Krishna, Diksha Chhabra, Ankur T. Patil, Madhu R. Kamble, Hemant A. Patil |
INTERSPEECH | 6 |
| 2018 | Auditory Filterbank Learning for Temporal Modulation Features in Replay Spoof Speech Detection
Hardik B. Sailor, Madhu R. Kamble, Hemant A. Patil |
INTERSPEECH | 3 |
| 2018 | Auditory Filterbank Learning Using ConvRBM for Infant Cry Classification
Hardik B. Sailor, Hemant A. Patil |
INTERSPEECH | 2 |
| 2018 | Unsupervised Vocal Tract Length Warped Posterior Features for Non-Parallel Voice Conversion
Nirmesh J. Shah, Maulik C. Madhavi, Hemant A. Patil |
INTERSPEECH | 3 |
| 2018 | Effectiveness of Dynamic Features in INCA and Temporal Context-INCA
Nirmesh J. Shah, Hemant A. Patil |
INTERSPEECH | 2 |
| 2018 | Effectiveness of Generative Adversarial Network for Non-Audible Murmur-to-Whisper Speech Conversion
Neil Shah, Nirmesh J. Shah, Hemant A. Patil |
INTERSPEECH | 3 |
| 2018 | Novel Linear Frequency Residual Cepstral Features for Replay Attack Detection
Hemlata Tak, Hemant A. Patil |
INTERSPEECH | 2 |
| 2018 | Novel Empirical Mode Decomposition Cepstral Features for Replay Spoof Detection
Prasad Tapkir, Hemant A. Patil |
INTERSPEECH | 2 |
| 2018 | Design of mixture of GMMs for Query-by-Example Spoken Term Detection
Maulik C. Madhavi, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2018 | Combining evidences from magnitude and phase information using VTEO for person recognition using humming
Hemant A. Patil, Maulik C. Madhavi |
Comput. Speech Lang. | 1 |
| 2017 | Quality assessment of voice converted speech using articulatory featuresabstractWe propose a novel application of the acoustic-to-articulatory inversion (AAI) towards a quality assessment of the voice converted speech. The ability of humans to speak effortlessly requires the coordinated movements of various articulators, muscles, etc. This effortless movement contributes towards a naturalness, intelligibility and speaker's identity (which is partially present in voice converted speech). Hence, during voice conversion (VC), the information related to the speech production is lost. In this paper, this loss is quantified for a male voice, by showing an increase in RMSE error (up to 12.7 % in tongue tip) for voice converted speech followed by showing a decrease in mutual information (I) (by 8.7 %). Similar results are obtained in the case of a female voice. This observation is extended by showing that the articulatory features can be used as an objective measure. The effectiveness of the proposed measure over MCD is illustrated by comparing their correlation with a Mean Opinion Score (MOS). Moreover, the preference score of MCD contradicted ABX test by 100 %, whereas the proposed measure supported ABX test by 45.8 % and 16.7 % in the case of female-to-male and male-to-female VC, respectively. Avni Rajpal, Nirmesh J. Shah, Mohammadi Zaki, Hemant A. Patil |
ICASSP | 4 |
| 2017 | Novel Amplitude Scaling method for bilinear frequency Warping-based Voice ConversionabstractIn Frequency Warping (FW)-based Voice Conversion (VC), the source spectrum is modified to match the frequency-axis of the target spectrum followed by an Amplitude Scaling (AS) to compensate the amplitude differences between the warped spectrum and the actual target spectrum. In this paper, we propose a novel AS technique which linearly transfers the amplitude of the frequency-warped spectrum using the knowledge of a Gaussian Mixture Model (GMM)-based converted spectrum without adding any spurious peaks. The novelty of the proposed approach lies in avoiding a perceptual impression of wrong formant location (due to perfect match assumption between the warped spectrum and the actual target spectrum in state-of-the-art AS method) leading to deterioration in converted voice quality. From subjective analysis, it is evident that the proposed system has been preferred 33.81% and 12.37% times more compared to the GMM and state-of-the-art AS method for voice quality, respectively. Similar to the quality conversion trade-offs observed by other studies in the literature, speaker identity conversion was 0.73% times more and 9.09% times less preferred over GMM and state-of-the-art AS-based method, respectively. Nirmesh J. Shah, Hemant A. Patil |
ICASSP | 2 |
| 2017 | Novel Variable Length Teager Energy Separation Based Instantaneous Frequency Features for Replay Detection
Hemant A. Patil, Madhu R. Kamble, Tanvina B. Patel, Meet H. Soni |
INTERSPEECH | 1 |
| 2017 | Unsupervised Filterbank Learning Using Convolutional Restricted Boltzmann Machine for Environmental Sound Classification
Hardik B. Sailor, Dharmesh M. Agrawal, Hemant A. Patil |
INTERSPEECH | 3 |
| 2017 | Unsupervised Representation Learning Using Convolutional Restricted Boltzmann Machine for Spoof Speech Detection
Hardik B. Sailor, Madhu R. Kamble, Hemant A. Patil |
INTERSPEECH | 3 |
| 2017 | Novel Shifted Real Spectrum for Exact Signal Reconstruction
Meet H. Soni, Rishabh Tak, Hemant A. Patil |
INTERSPEECH | 3 |
| 2017 | Partial matching and search space reduction for QbE-STD
Maulik C. Madhavi, Hemant A. Patil |
Comput. Speech Lang. | 2 |
| 2016 | Effectiveness of fundamental frequency (F0) and strength of excitation (SOE) for spoofed speech detectionabstractCurrent countermeasures used in spoof detectors (for speech synthesis (SS) and voice conversion (VC)) are generally phase-based (as vocoders in SS and VC systems lack phase-information). These approaches may possibly fail for non-vocoder or unit-selection-based spoofs. In this work, we explore excitation source-based features, i.e., fundamental frequency (F0) contour and strength of excitation (SoE) at the glottis as discriminative features using GMM-based classification system. We use F0and SoE1 estimated from speech signal through zero frequency (ZF) filtering method. Further, SoE2 is estimated from negative peaks of derivative of glottal flow waveform (dGFW) at glottal closure instants (GCIs). On the evaluation set of ASVspoof 2015 challenge database, the F0and SoEs features along with its dynamic variations achieve an Equal Error Rate (EER) of 12.41%. The source features are fused at score-level with MFCC and recently proposed cochlear filter cepstral coefficients and instantaneous frequency (CFCCIF) features. On fusion with MFCC (CFCCIF), the EER decreases from 4.08% to 3.26% (2.07% to 1.72%). The decrease in EER was evident on both known and unknown vocoder-based attacks. When MFCC, CFCCIF and source features are combined, the EER further decreased to 1.61%. Thus, source features captures complementary information than MFCC and CFCCIF used alone. Tanvina B. Patel, Hemant A. Patil |
ICASSP | 2 |
| 2016 | Analysis of natural and synthetic speech using Fujisaki modelabstractText-to-speech (TTS) synthesis systems are being advanced to achieve naturalness and intelligibility in synthetic speech. Unit selection-based synthesis (USS) and Hidden Markov Model-based text-to-speech synthesis systems (HTS) are recent techniques in this area. USS-based synthetic speech is known to be natural (due to concatenation of natural speech sound units). On the other hand, HTS-based speech is not as natural in perception as USS-based synthetic speech. Due to speech synthesis technologies, voice biometrics systems may face threats due to impostor attacks. Thus, it is important to study the differences that exist between natural and synthetic speech. In this context, we investigate the effectiveness of parameters of Fujisaki model for capturing Fundamental frequency (F0) contour variations in natural and synthetic speech. F0 contour of speech contains linguistic and non-linguistic information. Experimental results on several utterances from Gujarati (a low resourced language) demonstrate the effectiveness of phrase and accent components to analyze the difference between these two speeches. Variability in phrase and accent components suggests that synthetic speech differs in terms of prosodic information in excitation source as compared to natural speech. These findings may assist to distinguish these two speeches and provide an aid to alleviate impostor attacks. Tanvina B. Patel, Hemant A. Patil |
ICASSP | 2 |
| 2016 | Filterbank learning using Convolutional Restricted Boltzmann Machine for speech recognitionabstractConvolutional Restricted Boltzmann Machine (ConvRBM) as a model for speech signal is presented in this paper. We have developed ConvRBM with sampling from noisy rectified linear units (NReLUs). ConvRBM is trained in an unsupervised way to model speech signal of arbitrary lengths. Weights of the model can represent an auditory-like filterbank. Our proposed learned filterbank is also nonlinear with respect to center frequencies of subband filters similar to standard filterbanks (such as Mel, Bark, ERB, etc.). We have used our proposed model as a front-end to learn features and applied to speech recognition task. Performance of ConvRBM features is improved compared to MFCC with relative improvement of 5% on TIMIT test set and 7% on WSJ0 database for both Nov'92 test sets using GMM-HMM systems. With DNN-HMM systems, we achieved relative improvement of 3% on TIMIT test set over MFCC and Mel filterbank (FBANK). On WSJ0 Nov'92 test sets, we achieved relative improvement of 4-14% using ConvRBM features over MFCC features and 3.6-5.6% using ConvRBM filterbank over FBANK features. Hardik B. Sailor, Hemant A. Patil |
ICASSP | 2 |
| 2016 | Novel Nonlinear Prediction Based Features for Spoofed Speech Detection
Himanshu N. Bhavsar, Tanvina B. Patel, Hemant A. Patil |
INTERSPEECH | 3 |
| 2016 | Native Language Identification Using Spectral and Source-Based Features
Avni Rajpal, Tanvina B. Patel, Hardik B. Sailor, Maulik C. Madhavi, Hemant A. Patil, Hiroya Fujisaki |
INTERSPEECH | 5 |
| 2016 | Unsupervised Deep Auditory Model Using Stack of Convolutional RBMs for Speech Recognition
Hardik B. Sailor, Hemant A. Patil |
INTERSPEECH | 2 |
| 2016 | Novel Subband Autoencoder Features for Non-Intrusive Quality Assessment of Noise Suppressed Speech
Meet H. Soni, Hemant A. Patil |
INTERSPEECH | 2 |
| 2016 | Novel Subband Autoencoder Features for Detection of Spoofed Speech
Meet H. Soni, Tanvina B. Patel, Hemant A. Patil |
INTERSPEECH | 3 |
| 2016 | Novel Unsupervised Auditory Filterbank Learning Using Convolutional RBM for Speech RecognitionabstractTo learn auditory filterbanks, recently, we have proposed an unsupervised learning model based on convolutional restricted Boltzmann machine (RBM) with rectified linear units. In this paper, theory, training algorithm of our proposed model, and detailed analysis of learned filterbank are being presented. Learning of the model with different databases shows that the model is able to learn cochlear-like impulse responses that are localized in frequency-domain. An auditory-like scale obtained from filterbanks learned from clean and noisy datasets resembles the Mel scale, which is known to mimic perceptually relevant aspect of speech. We have experimented with both cepstral (denoted as ConvRBM-CC) as well as filterbank features (denoted as ConvRBM-BANK). On large vocabulary continuous speech recognition task, we achieved relative improvement of 7.21-17.8% in word error rate (WER) compared to Mel frequency cepstral coefficient (MFCC) features and 1.35-6.82% compared to Mel filterbank (FBANK) features. On AURORA 4 multicondition training database, the relative improvement in WER by 4.8-13.65% was achieved using a Hybrid Deep Neural Network-Hidden Markov Model (DNN-HMM) system with ConvRBM-CC features. Using ConvRBM-BANK features, we achieve absolute reduction of 1.25-3.85% in WER on AURORA 4 test sets compared to FBANK features. A context-dependent DNN-HMM system further improves performance with a relative improvement of 3.6-4.6% on an average for bigram 5k and tri-gram 5k language models. Hence, our proposed learned filterbank performs better than traditional MFCC and Mel-filterbank features for both clean and multicondition automatic speech recognition (ASR) tasks. A system combination of ConvRBM-BANK and FBANK features further improve performance in all ASR tasks. Cross-domain experiments where subband filters trained on one database are used for the ASR task of another database show that model learns generalized representations of speech signals. Hardik B. Sailor, Hemant A. Patil |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A novel filtering based approach for epoch extractionabstractIn this paper, we propose a novel algorithm which uses simple lowpass filtering as pre-processing for detection of epochs. Lowpass filtering with an appropriate cut-off frequency removes the effect of vocal tract characteristics as formants lie in relatively higher frequency regions. The method is evaluated on entire CMU-ARCTIC database consisting of the electroglottograph (EGG) signals. Noise robustness of the proposed algorithm is evaluated in the presence of additive white noise with various SNR levels. Experimental results show that lowpass filtering make the proposed algorithm noise robust. The method gives comparable or better results with the two state-of-the-art methods, viz., ZFR and SEDREAMS (which require apriori knowledge of the pitch period). In addition, the proposed method shows an improvement in identification accuracy. Pramod B. Bachhav, Hemant A. Patil, Tanvina B. Patel |
ICASSP | 2 |
| 2015 | Combining evidences from mel cepstral, cochlear filter cepstral and instantaneous frequency features for detection of natural vs. spoofed speechabstractSpeech synthesis and voice conversion techniques can pose threats to current speaker verification (SV) systems. For this purpose, it is essential to develop front end systems that are able to distinguish human speech vs. spoofed speech (synthesized or voice converted). In this paper, for the ASVspoof 2015 challenge, we propose a detector based on combination of cochlear filter cepstral coefficients (CFCC) and change in instantaneous frequency (IF), (i.e., CFCCIF) to detect natural vs. spoofed speech. The CFCCIF features were extracted at frame-level and Gaussian mixture model (GMM)based classification system was used. On the development set, the proposed features (i.e., CFCCIF) after fusion with Mel frequency cepstral coefficients (MFCC) features achieved an EER of 1.52 %, which is a significant reduction from MFCC (3.26 %) and CFCCIF (2.29 %) alone using 12-D static features. The EER further decreases to 0.89 % and 0.83 % for delta and delta-delta features, respectively. Experimental results on evaluation set show that fusion of MFCC and CFCCIF works relatively well with an EER of 0.41 % for known attacks and 2.013 % EER for unknown attacks. On an average, fusion of MFCC and CFCCIF features provided relatively best EER of 1.211 % for the challenge. Tanvina B. Patel, Hemant A. Patil |
INTERSPEECH | 2 |
| 2014 | Effectiveness of PLP-based phonetic segmentation for speech synthesisabstractIn this paper, use of Viterbi-based algorithm and spectral transition measure (STM)-based algorithm for the task of speech data labeling is being attempted. In the STM framework, we propose use of several spectral features such as recently proposed cochlear filter cepstral coefficients (CFCC), perceptual linear prediction cepstral coefficients (PLPCC) and RelAtive SpecTrAl (RASTA)-based PLPCC in addition to Mel frequency cepstral coefficients (MFCC) for phonetic segmentation task. To evaluate effectiveness of these segmentation algorithms, we require manual accurate phoneme-level labeled data which is not available for low resourced languages such as Gujarati (one of the official languages of India). In order to measure effectiveness of various segmentation algorithms, HMM-based speech synthesis system (HTS) for Gujarati has been built. From the subjective and objective evaluations, it is observed that Viterbi-based and STM with PLPCC-based segmentation algorithms work better than other algorithms. Nirmesh J. Shah, Bhavik B. Vachhani, Hardik B. Sailor, Hemant A. Patil |
ICASSP | 4 |
| 2014 | Chaotic mixed excitation source for speech synthesis
Hemant A. Patil, Tanvina B. Patel |
INTERSPEECH | 1 |
| 2013 | Nonlinear prediction of speech signal using volterra-wiener series
Hemant A. Patil, Tanvina B. Patel |
INTERSPEECH | 1 |
| 2012 | A comparison of waveform fractal dimension techniques for voice pathology classificationabstractIn this paper, an attempt is made to compare and analyze the various waveform fractal dimension techniques for voice pathology classification. Three methods of estimating the fractal dimension directly from the time-domain waveform have been compared. The methods used are Katz algorithm, Higuchi algorithm and the Hurst exponent calculated using the rescaled range (R/S) analysis. Furthermore, the effects of the window size, the base waveform used and score-level fusion with Mel frequency cepstral coefficients (MFCC) has also been evaluated. The features have been extracted from two different base waveforms, the speech signal and the Teager energy operator (TEO) phase of the speech signal. Experiments have been carried out on a subset of the Massachusetts Eye and Ear Infirmary (MEEI) database and classifier used is a 2ndorder polynomial classifier. A classification accuracy of 97.54 %was achieved on score-level fusion, an increase in performance by about 2 % as compared to MFCC alone. Pallavi N. Baljekar, Hemant A. Patil |
ICASSP | 2 |
| 2011 | Novel VTEO Based Mel Cepstral Features for Classification of Normal and Pathological Voices
Hemant A. Patil, Pallavi N. Baljekar |
INTERSPEECH | 1 |
| 2011 | Combining Evidence from Spectral and Source-Like Features for Person Recognition from Humming
Hemant A. Patil, Maulik C. Madhavi, Keshab K. Parhi |
INTERSPEECH | 1 |
| 2010 | Novel Variable length Teager Energy Based features for person recognition from their humabstractMost of the state-of-the-art voice biometrics systems use the natural speech signal (either read speech or spontaneous or contextual speech) from the subjects. In this paper, an attempt is made to identify speakers from their hum. A new feature set, viz., Variable length Teager Energy Based Mel Frequency Cepstral Coefficients (VTMFCC) is proposed for this problem. Experiments have been carried out for person identification and verification task using Linear Prediction Cepstral Coefficients (LPCC) and Mel Frequency Cepstral Coefficients (MFCC) with polynomial classifier of 2ndorder approximation. It is shown that the speaker identification rate for proposed feature set outperforms LPCC by 13.6% and is competitive over baseline MFCC. For speaker verification, a reduction in equal error rate (EER) by 1.73% is achieved when a score-level fusion system is employed by combining evidence from MFCC and VTMFCC. Hemant A. Patil, Keshab K. Parhi |
ICASSP | 1 |
| 2008 | On the development of variable length Teager energy operator (VTEO)abstractTeager Energy Operator (TEO) proposed by Kaiser and Teager is based on a definition of energy required to generate the signal. TEO gives us the running estimate of energy as a function of amplitude and instantaneous frequency content of the signal. However, it considers three consecutive samples to calculate the energy estimate. In this paper, we suggests an alternative and generalized approach to TEO to calculate the instantaneous estimate of the energy where not only consecutive but other distant samples can also be incorporated in the calculation of running estimate of the energy and the number of samples taken to calculate energy can also be increased depending on our signal properties to better capture the energy content variations in the speech signal. Vikrant Tomar, Hemant A. Patil |
INTERSPEECH | 2 |
| 2004 | The Teager Energy Based Features for Identification of Identical Twins in Multi-lingual Environment
Hemant A. Patil, Tapan Kumar Basu |
ICONIP | 1 |