VLDB 2026 Research / reviewers in the wild / expert
Syed Shahnawazuddin
dblp:145/1303 · also S. Shahnawazuddin
· DBLP profile ↗
31ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0002-3916-9693ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring LoRA variants to adapt whisper models for robust recognition of children's speech
Kancharana Manideep Bharadwaj, Syed Azim, Nagaraj Adiga, Syed Shahnawazuddin |
Speech Commun. | 5 |
| 2025 | On Enhancing the Performance of Children's ASR Task in Limited Data Scenario
Shambhavi, Syed Shahnawazuddin |
INTERSPEECH | 3 |
| 2025 | Advancing low-light image enhancement: an entropy-driven approach with reduced defects
Syed Shahnawazuddin |
Multim. Tools Appl. | 3 |
| 2025 | Detection of fricative and vowels in speech signals
Syed Shahnawazuddin |
Multim. Tools Appl. | 2 |
| 2024 | Combined approach to dysarthric speaker verification using data augmentation and feature fusion
Shinimol Salim, Syed Shahnawazuddin, Waquar Ahmad |
Speech Commun. | 2 |
| 2024 | Effect of Modeling Glottal Activity Parameters on Zero-Shot Children's ASRabstractThe primary objective of this study is to enhance the recognition performance ofzero-shot children'sautomatic speech recognition (ASR) task. In such a setup, statistical models are trained on adults' speech while the test data is from child speakers. Due to the stark differences in the speech data from adult and children, poorer recognition performances are obtained. To mitigate the ill-effects of acoustic mismatch between the training and the test data, we have first resorted to data augmentation wherein the attributes of adults' training speech were modified in such way that those become acoustically similar to children's speech. Data augmentation helped us build a better baseline ASR system. In order to further enhance the recognition performance, we have studied the impact of concatenating glottal activity parameters with the Mel-frequency cepstral coefficients (MFCC). The explored glottal activity parameters are strength of excitation (SoE), normalized auto-correlation peak strength (NAPS), higher-order statistics (HOS), jitter and shimmer. However, shimmer was found to be unsuitable for thezero-shot children'sASR task. Therefore, only SoE, NAPS, HOS and jitter parameters were included into the training and this, in turn, yielded significantly improved recognition performance. A relative reduction in word error rate by$24.5\%$over the baseline was obtained on including glottal activity parameters. Similarly, the relative reduction in character error rate was noted to be$30\%$over the baseline. Finally, in order to further build confidence in the idea of modeling glottal activity parameters, we have extended the study tolow-resource children'sASR task employing end-to-end architecture. Even in this case, significant reductions in error rates were noted. Shambhavi, Syed Shahnawazuddin |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Automatic Speaker Verification System for Dysarthria Patients
Shinimol Salim, Syed Shahnawazuddin, Waquar Ahmad |
INTERSPEECH | 2 |
| 2021 | A fused contextual color image thresholding using cuttlefish algorithm
Ashish Kumar Bhandari, Kusuma Rahul, Syed Shahnawazuddin |
Neural Comput. Appl. | 3 |
| 2020 | In-Domain and Out-of-Domain Data Augmentation to Improve Children's Speaker Verification System in Limited Data ScenarioabstractIn this paper, we present our efforts towards developing a robust automatic speaker verification (ASV) system for children when the domain-specific data is limited. For that purpose, we have studied the effect of in-domain and out-of-domain data augmentation. Several different combinations of data augmentation are studied in this work. Speed and pitch perturbation of children's speech are employed for synthetically creating in-domain data to be used for augmentation. For out-of-domain data augmentation, on the other hand, adults' speech is pooled together with children's speech. At the same time, voice conversion (VC) is also applied on adults' speech to alter the acoustic attributes. VC of adults' speech makes it perceptually similar to that of children's speech. The converted adults' data is then used for augmentation. The ASV systems developed in this study employ x-vectors derived using a time-delay deep neural network. In addition to that, probabilistic linear discriminant analysis is used for scoring the performance. The explored methods of data augmentation are noted to reduce the equal error rate as well as minimum decision cost function by a large margin. Syed Shahnawazuddin, Waquar Ahmad, Nagaraj Adiga |
ICASSP | 1 |
| 2020 | A Noise Robust Technique for Detecting Vowels in Speech Signals
Syed Shahnawazuddin, Waquar Ahmad |
INTERSPEECH | 2 |
| 2020 | Voice Conversion Based Data Augmentation to Improve Children's Speech Recognition in Limited Data Scenario
Syed Shahnawazuddin, Nagaraj Adiga, Aayushi Poddar, Waquar Ahmad |
INTERSPEECH | 1 |
| 2020 | Creating speaker independent ASR system through prosody modification based data augmentation
Syed Shahnawazuddin, Nagaraj Adiga, Hemant Kumar Kathania, B. Tarun Sai |
Pattern Recognit. Lett. | 1 |
| 2020 | A Novel Fuzzy Clustering-Based Histogram Model for Image Contrast EnhancementabstractHistogram equalization is a famous method for enhancing the contrast and image features. However, in few cases, it causes the overenhancement, and hence demolishes the natural display of the image. Therefore, in this article, a new fuzzy clustering based subhistogram scheme using discrete cosine transform (DCT) for contrast enhancement has been proposed. For preserving the distinctive appearance of the image, histogram division and separate histogram equalization is done on each subhistogram. The way of dividing histogram and calculating the numbers of parts for histogram division are the major problems which directly affects the quality of the output image. The proposed fuzzy-DCT scheme includes automatic calculation of a number of parts in which histogram is divided. Histogram division has done on the basis of density function and histogram separation is computed in such a way that each main peak can be divided in a different segment. The proposed scheme consists of four stages. The first stage includes the automatic calculation of number of clusters for image brightness levels. The second stage includes clustering of brightness levels by the fuzzy c-means clustering method and utilizing the given transfer function of histogram equalization. In the third stage, contrast enhancement is computed on each individual cluster separately. In the final stage, DCT is employed on the resulting image of the third step for better contrast and brightness preservation. The simulation results of the proposed scheme reveal not only clearer features along with a contrast enhancement, but also remarkably more natural look in the images. Ashish Kumar Bhandari, Syed Shahnawazuddin, Ayur Kumar Meena |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | Role of Prosodic Features on Children's Speech RecognitionabstractIn this paper, we have explored the role of combining prosodic variables with the existing acoustic features in the context of children's speech recognition using acoustic models trained on adults' speech. The explored acoustic features are Mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction cepstral coefficients (PLPCC) while the considered prosodic variables are loudness, voice-intensity and voice-probability. An analysis presented in this paper shows that, given that the textual content remains the same, the considered prosodic variables exhibit very similar contours for adults' and children's speech. At the same time, the contours differ a lot when the context is different. Consequently, inclusion of prosodic information reduces the inter-speaker differences and increases the class discrimination. This subsequently improves the recognition performance. Further improvements are obtained by projecting the feature vectors obtained by combining the two features to a lower-dimensional subspace. The same has been experimentally verified in this study for mismatched speech recognition using deep neural network (DNN) based system. On combining MFCC (PLPCC) and prosodic features, a relative improvement of 16% (14%) is noted on decoding children's speech using adult data trained DNN models. Hemant Kumar Kathania, Syed Shahnawazuddin, Nagaraj Adiga, Waquar Ahmad |
ICASSP | 2 |
| 2018 | Spectral Smoothing by Variationalmode Decomposition and its Effect on Noise and Pitch Robustness of ASR SystemabstractA novel front-end speech parameterization technique that is robust towards ambient noise and pitch variations is proposed in this paper. In the proposed technique, the short-time magnitude spectrum obtained by discrete Fourier transform is first decomposed in several components using variational mode decomposition (VMD). For sufficiently smoothing the spectrum, the higher-order components are discarded. The smoothed spectrum is then obtained by reconstructing the spectrum using the first-two modes only. The Mel-frequency cepstral coefficients computed using the VMD-based smoothed spectra are observed to be affected less by ambient noise and pitch variations. To validate the same, an automatic speech recognition system is developed on clean speech from adult speakers and evaluated under noisy test conditions. Furthermore, experimental evaluations are also performed on another test set which consists of speech data from children to simulate large pitch differences. The experimental evaluations as well as signal domain analyses presented in this paper support these claims. Ishwar Chandra Yadav, Syed Shahnawazuddin, D. Govind 0001, Gayadhar Pradhan |
ICASSP | 2 |
| 2018 | Exploring Sparse Representation for Improved Online Handwriting RecognitionabstractThis work studies the sparse representation based classification (SRC) framework for online handwriting recognition (HR) task. In this framework, first, an exemplar dictionary is created using the training samples from each of the classes in the chosen task. Subsequently, the test samples are sparse coded over exemplar dictionary for classification. In sparse coding, both l_0 - and l_1 -norm based greedy algorithms are studied. Further, for reducing the computational cost of the SRC-based HR approach, the learned exemplar dictionary has also been explored. The proposed SRC-based approach is demonstrated for character and limited vocabulary word recognition task and evaluated on three different corpora: the Assamese digit database, the UNIPEN English character database and the UNIPEN ICROW-03 English word database. The experimental results are promising over the reported works on these databases employing the hidden Markov model or the support vector machine. Subhasis Mandal, Syed Shahnawazuddin, Rohit Sinha 0003, S. R. Mahadeva Prasanna, Suresh Sundaram 0001 |
ICFHR | 2 |
| 2018 | Enhancement of Noisy Speech Signal by Non-Local Means Estimation of Variational Mode Functions
Nagapuri Srinivas, Gayadhar Pradhan, Syed Shahnawazuddin |
INTERSPEECH | 3 |
| 2018 | Non-Uniform Spectral Smoothing for Robust Children's Speech Recognition
Ishwar Chandra Yadav, Syed Shahnawazuddin, Gayadhar Pradhan |
INTERSPEECH | 3 |
| 2018 | Assessment of pitch-adaptive front-end signal processing for children's speech recognitionabstractOn account of large acoustic mismatch, automatic speech recognition (ASR) systems trained using adults’ speech data yield poor recognition performance when evaluated on children’s speech data. Despite the use of common speaker normalization techniques like feature-space maximum likelihood regression (fMLLR) and vocal tract length normalization (VTLN), a significant gap remains between the recognition rates for matched and mismatched testing. Our earlier works have already highlighted the sensitivity of salient front-end features including the popular Mel-frequency cepstral coefficient (MFCC) to gross pitch variation across adult and child speakers. Motivated by that, in this work, we explore pitch-adaptive front-end signal processing in deriving the MFCC features to reduce the sensitivity to pitch variation. For this purpose, first an existing vocoder approach known as STRAIGHT spectral analysis is employed for obtaining the smoothed spectrum devoid of pitch harmonics. Secondly, a much simpler spectrum smoothing approach exploiting pitch adaptive-liferting is also presented. The proposed approach is noted to be less sensitive to errors in the pitch estimation than the STRAIGHT-based approach. Both these approaches result in significant improvements for children’s mismatch ASR. The effectiveness of the proposed adaptive-liftering-based approach is also demonstrated in the context of acoustic modeling paradigms based on the subspace Gaussian mixture model (SGMM) and the deep neural network (DNN). Further, it has been shown that the effectiveness of existing speaker normalization techniques remain intact even with the use of proposed pitch-adaptive MFCCs, thus leading to additional gains. Rohit Sinha 0003, Syed Shahnawazuddin |
Comput. Speech Lang. | 2 |
| 2018 | Improving children's mismatched ASR using structured low-rank feature projection
Syed Shahnawazuddin, Hemant Kumar Kathania, Abhishek Dey, Rohit Sinha 0003 |
Speech Commun. | 1 |
| 2017 | Enhancing noise and pitch robustness of children's ASRabstractIt is well known that, when noisy speech is transcribed using automatic speech recognition (ASR) systems trained on clean data, a highly degraded recognition performance is obtained. The problemgets further aggravatedwhen the targeted group happens to be child speakers. For children's speech, the acoustic correlates such as pitch and formant frequency vary significantly with age. This makes the recognition of children's speech very challenging. In this paper, we have explored the ways to enhance the noise robustness of ASR systems for children's speech. Towards addressing the same, recently developed front-end acoustic features based on spectral moments (SMAC) are explored. The SMAC features are reported to be more noise robust than the conventional features like the mel-frequency cepsatral coefficients. At the same time, the SMAC features are also noted to be sensitive to the variations in the pitch. To reduce the pitch sensitivity, a spectral smoothing approach based on adaptive-liftering is proposed. Spectral smoothening prior to the computation of spectral moments results in a significant improvement in the robustness to pitch without affecting the noise immunity. To further enhance noise robustness, a foreground speech segmentation and enhancement module is also included in the proposed front-end speech parameterization technique. Syed Shahnawazuddin, K. T. Deepak, Gayadhar Pradhan, Rohit Sinha 0003 |
ICASSP | 1 |
| 2017 | Improving Children's Speech Recognition Through Explicit Pitch Scaling Based on Iterative Spectrogram Inversion
Waquar Ahmad, Syed Shahnawazuddin, Hemant Kumar Kathania, Gayadhar Pradhan, Arun B. Samaddar |
INTERSPEECH | 2 |
| 2017 | Non-Local Estimation of Speech Signal for Vowel Onset Point Detection in Varied Environments
Syed Shahnawazuddin, Gayadhar Pradhan |
INTERSPEECH | 2 |
| 2017 | Excitation Source Features for Improving the Detection of Vowel Onset and Offset Points in a Speech Sequence
Gayadhar Pradhan, Syed Shahnawazuddin |
INTERSPEECH | 3 |
| 2017 | Sparse coding over redundant dictionaries for fast adaptation of speech recognition system
Syed Shahnawazuddin, Rohit Sinha 0003 |
Comput. Speech Lang. | 1 |
| 2017 | Effect of Prosody Modification on Children's ASRabstractTranscribing children's speech using acoustic models trained on adults' speech is very challenging. In such conditions, a highly degraded recognition performance is reported due to large mismatch in the acoustic/linguistic attributes of the training and test data. The differences in pitch (or fundamental frequency) between the two groups of speakers is one among several mismatch factors. Another important mismatch factor is the difference in speaking rates. To overcome these two sources of mismatch, prosody modification is explored in this letter. Prosody modification is done by using glottal closure instants (GCIs) as anchoring points. The GCIs, in turn, are determined using zero-frequency filtering (ZFF). The ZFF-GCI-based prosody modification is fast and results in highly accurate scaling of pitch and speaking rate. The experimental evaluations studying the effect of prosody modification resulted in a relative improvement of 50% over the baseline. Syed Shahnawazuddin, Nagaraj Adiga, Hemant Kumar Kathania |
IEEE Signal Process. Lett. | 1 |
| 2017 | Pitch-Normalized Acoustic Features for Robust Children's Speech RecognitionabstractIn this letter, the effectiveness of recently reported SMAC (Spectral Moment time-frequency distribution Augmented by low-order Cepstral) features has been evaluated for robust automatic speech recognition (ASR). The SMAC features consist of normalized first central spectral moments appended with low-order cepstral coefficients. These features have been designed for achieving robustness to both additive noise and the pitch variations. We have explored the SMAC features in severe pitch mismatch ASR task, i.e., decoding of children's speech on adults' speech trained ASR system. In those tasks, the SMAC features are still observed to be sensitive to pitch variations. Toward addressing the same, a simple spectral smoothening approach employing adaptive-cepstral truncation is explored prior to the computation of spectral moments. With the proposed modification, the SMAC features are noted to achieve enhanced pitch robustness without affecting their noise immunity. Furthermore, the effectiveness of the proposed features is explored in three dominant acoustic modeling paradigms and varying data conditions. In all the cases, the proposed features are observed to significantly outperform the existing ones. Syed Shahnawazuddin, Rohit Sinha 0003, Gayadhar Pradhan |
IEEE Signal Process. Lett. | 1 |
| 2016 | Pitch-Adaptive Front-End Features for Robust Children's ASR
Syed Shahnawazuddin, Abhishek Dey, Rohit Sinha 0003 |
INTERSPEECH | 1 |
| 2015 | Low-memory fast on-line adaptation for acoustically mismatched children's speech recognition
Syed Shahnawazuddin, Rohit Sinha 0003 |
INTERSPEECH | 1 |
| 2014 | A low complexity model adaptation approach involving sparse coding over multiple dictionaries
Syed Shahnawazuddin, Rohit Sinha 0003 |
INTERSPEECH | 1 |
| 2014 | Improved Bases Selection in Acoustic Model Interpolation for Fast On-Line AdaptationabstractThis work presents a novel bases selection approach for acoustic model interpolation based fast on-line adaptation. The proposed approach employs a correlation based similarity measure in the supervector domain (derived by concatenating the Gaussian mean parameters of the adapted models) for the selection of bases. This approach is found to greatly reduce the computational complexity in comparison to the Viterbi-alignment based bases search. Moreover, the proposed approach employs joint representation along with orthogonalization for the dynamic selection of bases. Consequently, the selected bases result in a much balanced coverage of phonetic contexts in the synthesized adapted model. The proposed technique is found to result in improved performance for all three modes of adaptation viz. the utterance-specific, the incremental and the batch modes. For utterance-specific mode, it achieves a relative improvement of 10.2% over baseline with only 3 to 5 seconds of adaptation data. Syed Shahnawazuddin, Rohit Sinha 0003 |
IEEE Signal Process. Lett. | 1 |