M. Ali Basha Shaik

dblp:98/9877 · also Mahaboob Ali Basha Shaik · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0002-5520-0830ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Elucidating Clock-drift Using Real-world Audios In Wireless Mode For Time-offset Insensitive End-to-End Asynchronous Acoustic Echo Cancellation
Premanand Nayak, M. Ali Basha Shaik
INTERSPEECH2
2024 Multi-mic Echo Cancellation Coalesced with Beamforming for Real World Adverse Acoustic Conditions
Premanand Nayak, Kamini Sabu, M. Ali Basha Shaik
INTERSPEECH3
2023 Dynamic Encoder RNN for Online Voice Activity Detection in Adverse Noise Conditions
Prithvi R. R. Gudepu, Jayesh M. Koroth, Kamini Sabu, M. Ali Basha Shaik
INTERSPEECH4
2023 Language Identification Networks for Multilingual Everyday Recordings
Kiran Praveen, Balaji Radhakrishnan, Kamini Sabu, M. Ali Basha Shaik
INTERSPEECH5
2022 Transformer Networks for Non-Intrusive Speech Quality Prediction
abstract
This paper presents the details of our speech quality prediction system submitted to the Conferencing Speech-2022 challenge.The challenge involved the task of non-intrusive speech quality assessment intended for online conferencing applications.We propose two approaches for speech quality prediction in this work.The first approach uses a combination of deep convolutional neural network (CNN) and LSTM neural network with Kullback-Leibler (KL) loss function and cross entropy (CE) loss function for estimating the mean opinion scores (MOS).Our second approach uses transformer based encoder network before applying attention pooling.We observe that our proposed second method gives significant improvements compared to our first method as well as on the baselines provided by the challenge organizers with respect to Pearson Correlation Coefficient (PCC) and Spearman Rank Correlation Coefficient (SRCC) along with reductions in root mean square error (RMSE).The model is also seen to generalize for unseen data resources on the evaluation dataset.
M. K. Jayesh, Mukesh Sharma, Praneeth Vonteddu, M. Ali Basha Shaik, Sriram Ganapathy
INTERSPEECH4
2021 SRIB-LEAP Submission to Far-Field Multi-Channel Speech Enhancement Challenge for Video Conferencing
abstract
This paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from microphone arrays in a video conferencing room. We propose a two stage method involving a beamformer followed by single channel enhancement. For the beamformer, we incorporated self-attention mechanism as inter-channel processing layer in the filter-and-sum network (FaSNet), an end-to-end time-domain beamforming system. The single channel speech enhancement is done in log spectral domain using convolution neural network (CNN)-long short term memory (LSTM) based architecture. We achieved improvements in objective quality metrics - perceptual evaluation of speech quality (PESQ) of 0.5 on the noisy data. On subjective quality evaluation, the proposed approach improved the mean opinion score (MOS) by an absolute measure of 0.9 over the noisy audio.
R. G. Prithvi Raj, M. K. Jayesh, Anurenjan Purushothaman, Sriram Ganapathy, M. Ali Basha Shaik
Interspeech6
2020 Whisper Augmented End-to-End/Hybrid Speech Recognition System - CycleGAN Approach
Prithvi R. R. Gudepu, Gowtham P. Vadisetti, Abhishek Niranjan, Kinnera Saranu, Raghava Sarma, M. Ali Basha Shaik, Periyasamy Paramasivam
INTERSPEECH6
2019 Improving Grapheme-to-Phoneme Conversion by Investigating Copying Mechanism in Recurrent Architectures
abstract
Attention driven encoder-decoder architectures have become highly successful in various sequence-to-sequence learning tasks. We propose copy-augmented Bi-directional Long Short-Term Memory based Encoder-Decoder architecture for the Grapheme-to-Phoneme conversion. In Grapheme-to-Phoneme task, a number of character units in words possess high degree of similarity with some phoneme unit(s). Thus, we make an attempt to capture this characteristic using copy-augmented architecture. Our proposed model automatically learns to generate phoneme sequences during inference by copying source token embeddings to the decoder's output in a controlled manner. To our knowledge, this is the first time the copy-augmentation is being investigated for Grapheme-to-Phoneme conversion task. We validate our experiments over accented and non-accented publicly available CMU-Dict datasets and achieve State-of-The-Art performances in terms of both phoneme and word error rates. Further, we verify the applicability of our proposed approach on Hindi Lexicon and show that our model outperforms all recent State-of-The-Art results.
Abhishek Niranjan, M. Ali Basha Shaik
ASRU2
2015 Improved strategies for a zero oov rate LVCSR system
abstract
In this work, multiple hierarchical language modeling strategies for a zero OOV rate large vocabulary continuous speech recognition system are investigated. In our previously proposed hierarchical approach, a full-word language model and a context independent character-level LM (CLM) are directly used during search. The novelty of this work is to jointly model the character-level prior and the pronunciation probabilities, to introduce across-word context into the characterlevel LM, and to properly normalize the character-level LM using prefix-tree based normalization for the hierarchical approach. Significant reductions in-terms of word error rates (WER) on the best full-word Quaero Polish LVCSR system are reported.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Stefan Hahn, Ralf Schlüter, Hermann Ney
ICASSP1
2015 Investigation of Segmental Conditional Random Fields for large vocabulary handwriting recognition
abstract
Multiple types of models are used in handwriting recognition and can be broadly categorized into generative and discriminative models. Gaussian Hidden Markov Models are used successfully in most of the systems. Discriminative training can be applied to these models to improve them further. Alternatively, Segmental Conditional Random Fields have the advantage of being discriminative as well as segmental. The novelty of this work is the investigation of Segmental Conditional Random Fields for handwriting recognition. In addition, Multi-Layer Perceptrons and Long Short Term Memory Recurrent Neural Networks are compared for the observations generation in this framework. Various types of features are investigated in the segmental models for handwriting recognition. Furthermore, class-based language model features are proposed to extend this model. Visual features based on moments are extracted at a word level to make the model more robust. Experimental results on English handwriting show a relative reduction of 13.7% in terms of word error rate w.r.t. the baseline system. The proposed system also outperforms the Gaussian Hidden Markov Models trained discriminatively using the minimum phone error criterion by a relative reduction of 6.9% in terms of word error rate.
Mahdi Hamdani, M. Ali Basha Shaik, Patrick Doetsch, Hermann Ney
ICDAR2
2015 Improvements in RWTH LVCSR evaluation systems for Polish, Portuguese, English, urdu, and Arabic
abstract
In this work, Portuguese, Polish, English, Urdu, and Arabic automatic speech recognition evaluation systems developed by the RWTH Aachen University are presented. Our LVCSR systems focus on various domains like broadcast news, spontaneous speech, and podcasts. All these systems but Urdu are used for Euronews and Skynews evaluations as part of the EUBridge project. Our previously developed LVCSR systems were improved using different techniques for the aforementioned languages. Significant improvements are obtained using multilingual tandem and hybrid approaches, minimum phone error training, lexical adaptation, open vocabulary long short term memory language models, maximum entropy language models and confusion-network based system combination. Index Terms: LVCSR, LSTM, open-vocabulary, EU-Bridge
M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2014 RWTH LVCSR systems for quaero and EU-bridge: German, Polish, Spanish and Portuguese
abstract
In this paper, German, Polish, Spanish, and Portuguese large vocabulary continuous speech recognition (LVCSR) systems developed by the RWTH Aachen University are presented.All the above mentioned systems for the aforementioned languages are used for the Quaero and EU-Bridge project evaluations.The LVCSR systems developed for these competitive evaluations focus on various domains like broadcast news, podcasts and lecture domain.Transcription of the speech for these tasks is challenging due to huge variability in the acoustic conditions and a significant portion of audio data includes spontaneous speech.Good improvements are obtained using stateof-the-art multilingual bottleneck features, minimum phone error trained acoustic models, language model (LM) adaptation and confusion-network based system combination.In addition, an open vocabulary approach using morphemic units is investigated along with the LM adaptation for the German LVCSR.
M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2013 Morpheme level hierarchical pitman-yor class-based language models for LVCSR of morphologically rich languages
abstract
Performing large vocabulary continuous speech recognition (LVCSR) for morphologically rich languages is considered a challenging task.The morphological richness of such languages leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities.In this case, the use of morphemes has been shown to increase the lexical coverage and lower the LM perplexity.Another approach used to improve the LM probability estimates is to incorporate additional knowledge sources in the LM estimation process using classbased LMs (CLMs).Recently, the hierarchical Pitman-Yor LMs (HPYLMs) have shown superiority over the modified Kneser-Ney (MKN) smoothed N-gram LMs in terms of both perplexity (PPL) and word error rate (WER) on word-based LVCSR tasks.In this paper, hierarchical Pitman-Yor class-based LMs (HPY-CLMs) are combined with morpheme level language modeling.This enables the application of the proposed models on top of morpheme-based systems.Experiments are conducted on Arabic and German LVCSR tasks.Consistent performance improvements are obtained for all the available corpora compared to the conventional morpheme-based and class-based LMs.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2013 Feature-rich sub-lexical language models using a maximum entropy approach for German LVCSR
abstract
German is a morphologically rich language having a high degree of word inflections, derivations and compounding. This leads to high out-of-vocabulary (OOV) rates and poor language model (LM) probabilities in the large vocabulary continuous speech recognition (LVCSR) systems. One of the main challenges in the German LVCSR is the recognition of the OOV words. For this purpose, data-driven morphemes are used to provide higher lexical coverage. On the other hand, the probability estimates of a sub-lexical LM could be further improved using feature-rich LMs like maximum entropy (MaxEnt) and class-based LMs. In this work, for a sub-lexical level German LVCSR task, we investigate the use of the multiple morpheme level features as classes for building class-based LMs that are estimated using the state-of-the-art MaxEnt approach. Thus, the benefits of both the MaxEnt LMs and the traditional class-based LMs are effectively combined. Furthermore, we experiment the use of Maximum a-posteriori adaptation over the MaxEnt class-based LMs. We show consistent reductions in both the OOV recognition error rate and the word error rate (WER) on a German LVCSR task from the Quaero project, compared to the traditional class-based and theN -gram morpheme based LM. Index Terms: open-vocabulary, German LVCSR, features, maximum entropy, class-based
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2012 Morpheme Level Feature-based Language Models for German LVCSR
abstract
One of the challenges for Large Vocabulary Continuous Speech Recognition (LVCSR) of German is its complex morphology and high level of compounding.It leads to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probabilities.In such cases, building LMs on morpheme level can be considered a better choice.Thereby, higher lexical coverage and lower LM perplexities are achieved.On the other side, a successful approach to improve the LM probability estimation is to incorporate features of words using feature-based LMs.In this paper, we use features derived for morphemes as well as words.Thus, we combine the benefits of both morpheme level and feature rich modeling.We compare the performance of stream-based, class-based and factored LMs (FLMs).Relative reductions of around 1.5% in Word Error Rate (WER) are achieved compared to the best previous results obtained using FLMs.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2012 Investigation of Maximum Entropy Hybrid Language Models for Open Vocabulary German and Polish LVCSR
abstract
For languages like German and Polish, higher numbers of word inflections lead to high out-of-vocabulary (OOV) rates and high language model (LM) perplexities.Thus, one of the main challenges in large vocabulary continuous speech recognition (LVCSR) is recognizing an open vocabulary.In this paper, we investigate the use of mixed type of sub-word units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables, along with full-words.In addition, we investigate the suitability of hybrid mixed-unit Ngrams as features for Maximum Entropy LM along with adaptation.We achieve significant improvements in recognizing OOVs and word error rate reductions for German and Polish LVCSR compared to the conventional full-word approach and state-of-the-art N-gram mixed type hybrid LM.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2011 Using morpheme and syllable based sub-words for polish LVCSR
abstract
Polish is a synthetic language with a high morpheme-per-word ratio. It makes use of a high degree of inflection leading to high out-of-vocabulary (OOV) rates, and high Language Model (LM) perplexities. This poses a challenge for Large Vocabulary and Continuous Speech Recognition (LVCSR) systems. Here, the use of morpheme and syllable based units is investigated for building sub-lexical LMs. A different type of sub-lexical units is proposed based on combining morphemic or syllabic units with corresponding pronunciations. Thereby, a set of grapheme-phoneme pairs called graphones are used for building LMs. A relative reduction of 3.5% in Word Error Rate (WER) is obtained with respect to a traditional system based on full-words.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
ICASSP1
2011 Morpheme Based Factored Language Models for German LVCSR
abstract
German is a highly inflectional language, where a large number of words can be generated from the same root.It makes a liberal use of compounding leading to high Out-of-vocabulary (OOV) rates, and poor Language Model (LM) probability estimates.Therefore, the use of morphemes for language modeling is considered a better choice for Large Vocabulary Continuous Speech Recognition (LVCSR) than the full-words.Thereby, better lexical coverage and less LM perplexities are achieved.On the other side, the use of Factored Language Models (FLMs) is considered a successful approach that allows the integration of many information sources to get better LM probability estimates.In this paper, we try a combined methodology for language modeling where both morphological decomposition and factored language modeling are used in one model called morpheme based FLM.Finally, we obtain around 2.5% relative reduction in Word Error Rate (WER) with respect to a traditional full-words system.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
INTERSPEECH2
2011 Hybrid Language Models Using Mixed Types of Sub-Lexical Units for Open Vocabulary German LVCSR
abstract
German is a highly inflected language with a large number of words derived from the same root.It makes use of a high degree of word compounding leading to high Out-of-vocabulary (OOV) rates, and Language Model (LM) perplexities.For such languages the use of sub-lexical units for Large Vocabulary Continuous Speech Recognition (LVCSR) becomes a natural choice.In this paper, we investigate the use of mixed types of sub-lexical units in the same recognition lexicon.Namely, morphemic or syllabic units combined with pronunciations called graphones, normal graphemic morphemes or syllables along with full-words.This mixture of units is used for building hybrid LMs suitable for open vocabulary LVCSR where the system operates over an open, constantly changing vocabulary like in broadcast news, political debates, etc.A relative reduction of around 5.0% in Word Error Rate (WER) is obtained compared to a traditional full-words system.Moreover, around 40% of the OOVs are recognized.
M. Ali Basha Shaik, Amr El-Desoky Mousa, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2010 Sub-lexical language models for German LVCSR
abstract
One of the major difficulties related to German LVCSR is the rich morphology nature of German, leading to high out-of-vocabulary (OOV) rates, and high language model (LM) perplexities. Normally, compound words make up an essential fraction of the German vocabulary. Most compound OOVs are composed of frequent in-vocabulary words. Here, we investigate the use of sub-lexical LMs based on different approaches for word decomposition, namely supervised and unsupervised decomposition, as well as decomposition derived from grapheme-to-phoneme (G2P) conversion. In the later approach, we augment a normal word model with a set of grapheme-phoneme pairs called graphones used to model the OOV words. A novel approach is proposed to select the representative graphone sequences for OOVs based on unsupervised decomposition and word-pronunciation alignment. We obtain relative reductions in word error rate (WER) from 4.2% to 6.5% with respect to a comparable full-words system.
Amr El-Desoky Mousa, M. Ali Basha Shaik, Ralf Schlüter, Hermann Ney
SLT2