Pankaj Warule

dblp:325/8920 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8201-7663ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Fixed frequency range empirical wavelet transform based acoustic and entropy features for speech emotion recognition
Siba Prasad Mishra, Pankaj Warule
Speech Commun.2
2024 Speech Emotion Recognition with DNN and Combination of CNN-LSTM
abstract
Emotion is essential to all living things. Understanding emotion is challenging for everyone, but it can solve thousands of problems and save many lives if done correctly. Emotion is represented in gestures, facial expression, speech, etc. Recording of speech is easy with non-invasive nature of processing. As a result, throughout the past three decades, numerous researchers have been interested in emotion recognition through speech. In this study, We address the crucial task of speech emotion recognition (SER) by leveraging deep learning architectures, namely Convolutional Neural Networks (CNN) combined with Long Short-Term Memory (LSTM) layers and Deep Neural Networks (DNN). our approach utilises a diverse feature set, including Mel spectrogram, MFCCs, spectrogram, delta of MFCCs, chroma feature, and tonal centroid features extracted from the EmoDB dataset. Through extensive exper-imentation employing a 5-fold cross-validation methodology, we achieved notable classification accuracies of 77.38% with the CNN+LSTM model and 87.09% with the DNN model.
Ponna Dinesh, Siba Prasad Mishra, Pankaj Warule
TENCON3
2024 Audio Fingerprinting Using Clustering
abstract
In the vast landscape of digital audio, the need for robust and efficient methods of identifying and managing audio content has become increasingly imperative. Audio fingerprinting emerges as a powerful solution to this challenge, offering a sophisticated means of uniquely characterizing audio signals for identification and retrieval purposes. Audio fingerprinting is a technique where a unique signature of music file that serves as a compact representation of audio is listened to and stored in an archive. When there is playback of that audio file anywhere, that audio is recognized, and it can be matched against a database that contains all the artist and ownership information for that copyright. The fingerprint feature should be based on perceptual features that are invariant with respect to signal degradation. The applications of audio fingerprinting are diverse and impactful. From music recognition and content matching in streaming services to copyright enforcement and audio-based search engines, technology plays a pivotal role in reshaping how we interact with and manage audio data in the digital era.
Sarth Patel, Anirudh PK, Aditi Pandey, Siba Prasad Mishra, Pankaj Warule, Sunandita Debnath
TENCON5
2024 Adaptive Variational Mode Decomposition Based Parkinson Disease Detection with Entropy-Based Features
Ramesh Chandra Pola, Pankaj Warule, Siba Prasad Mishra, Sunandita Debnath
TENCON2
2024 Automatic Detection of Parkinson's Disease Using Continuous Wavelet Transform-Based Time-Frequency Domain Analysis of Speech Signals
Pankaj Warule, Siba Prasad Mishra
TENCON1
2023 Hilbert-Huang Transform-Based Time-Frequency Analysis of Speech Signals for the Identification of Common Cold
abstract
The current advancements in machine learning research pertaining to speech and health are highly interesting. One aspect of speech-processing research that is gaining popularity is the use of computational paralinguistic analysis to evaluate a variety of health conditions. In this study, we have used the Hilbert-Huang transform (HHT) for the time-frequency analysis of speech signals for the identification of the common cold. The HHT is a time-frequency transform that is adaptive and ideal for non-linear and non-stationary signals. The HHT is a combination of empirical mode decomposition (EMD) and the Hilbert transform (HT). The HHT gives the time-frequency representation (TFR) matrix of the speech signal. Then, the entropy of each frequency component in TFR is computed and used as a distinguishing feature between cold and healthy speech. The efficacy of the proposed methodology is evaluated on the URTIC dataset using a deep neural network. The proposed features achieve UARs of 65.66% and 65.26%, respectively, on the develop and test partitions. The results of the study demonstrate that the time-frequency entropy features extracted using the HHT are effective in distinguishing between cold and healthy speech.
Pankaj Warule, Siba Prasad Mishra
TENCON1
2023 Empirical Mode Decomposition Based Detection of Common Cold Using Speech Signal
abstract
This study investigates the discrimination between cold speech and healthy speech using features based on empirical mode decomposition (EMD). The EMD is employed to break down the signal into several intrinsic mode functions (IMFs). From each IMF, various statistical values like minimum, maximum, mean, standard deviation, first, second, and third quartiles, skewness, kurtosis, and energy of each IMF are extracted and used as a feature for distinguishing cold and healthy speech. The T-test examines the importance of EMD-based features for classifying cold speech. EMD-based feature performance is assessed using the deep neural network (DNN) classifier. The findings show that EMD-based features effectively discriminate between cold and healthy speech classes. Combining Mel-Frequency Cepstral Coefficients (MFCC) characteristics with EMD-based features improves the performance for identifying healthy and cold speech classes. On the URTIC database, the combination of MFCC and EMD-based features achieve a UAR of 66.92%.
Pankaj Warule, Siba Prasad Mishra
TENCON1
2023 Chirplet transform based time frequency analysis of speech signal for automated speech emotion recognition
Siba Prasad Mishra, Pankaj Warule
Speech Commun.2