Rao Muhammad Adeel Nawab

dblp:53/8427 · also Adeel Nawab, Rao M. A. Nawab · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
13since 2021 · last 2024
0000-0002-1765-8904ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2024 Mono-lingual text reuse detection for the Urdu language at lexical level
Ayesha Noreen, Iqra Muneer, Rao Muhammad Adeel Nawab
Eng. Appl. Artif. Intell.3
2023 Urdu Text Reuse Detection at Phrasal level using Sentence Transformer-based approach
Gull Mehak, Iqra Muneer, Rao Muhammad Adeel Nawab
Expert Syst. Appl.3
2023 Tran-Switch: A transfer learning approach for sentence level cross-genre author profiling on code-switched English-RomanUrdu Text
Muhammad Adnan Ashraf, Rao Muhammad Adeel Nawab, Feiping Nie 0001
Inf. Process. Manag.2
2023 UNLT: Urdu Natural Language Toolkit
abstract
Abstract This study describes a Natural Language Processing (NLP) toolkit, as the first contribution of a larger project, for an under-resourced language—Urdu. In previous studies, standard NLP toolkits have been developed for English and many other languages. There is also a dire need for standard text processing tools and methods for Urdu, despite it being widely spoken in different parts of the world with a large amount of digital text being readily available. This study presents the first version of the UNLT (Urdu Natural Language Toolkit) which contains three key text processing tools required for an Urdu NLP pipeline; word tokenizer, sentence tokenizer, and part-of-speech (POS) tagger. The UNLT word tokenizer employs a morpheme matching algorithm coupled with a state-of-the-art stochastic n -gram language model with back-off and smoothing characteristics for the space omission problem. The space insertion problem for compound words is tackled using a dictionary look-up technique. The UNLT sentence tokenizer is a combination of various machine learning, rule-based, regular-expressions, and dictionary look-up techniques. Finally, the UNLT POS taggers are based on Hidden Markov Model and Maximum Entropy-based stochastic techniques. In addition, we have developed large gold standard training and testing data sets to improve and evaluate the performance of new techniques for Urdu word tokenization, sentence tokenization, and POS tagging. For comparison purposes, we have compared the proposed approaches with several methods. Our proposed UNLT, the training and testing data sets, and supporting resources are all free and publicly available for academic use.
Jawad Shafi, Hafiz Rizwan Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson
Nat. Lang. Eng.3
2023 Urdu Short Paraphrase Detection at Sentence Level
abstract
Paraphrase detection systems uncover the relationship between two text fragments and classify them as paraphrased when they convey the same idea; otherwise non-paraphrased. Previously, the researchers have mainly focused on developing resources for the English language for paraphrase detection. There have been very few efforts for paraphrase detection in South Asian languages. However, no research has been conducted on sentence-level paraphrase detection in Urdu, a low-resourced language. It is mainly due to the unavailability of the corpora that focus on the sentence level. The available related studies on the Urdu language only focus on text reuse detection tasks at the passage and document levels. Therefore, this study aims to develop a large-scale manually annotated benchmark Urdu paraphrase detection corpus at the sentence level, based on real cases from journalism. The proposed Urdu Sentential Paraphrases (USP) corpus contains 4,900 sentences (2,941 paraphrased and 1,959 non-paraphrased), manually collected from the Urdu newspapers. Moreover, several techniques were proposed, developed, and compared as a secondary contribution, including Word Embedding (WE), Sentence Transformers (ST), and feature-fusion techniques. N-gram is treated as the baseline technique for our research. The experimental results indicate that our proposed feature-fusion technique is the most suitable for the Urdu paraphrase detection task. Furthermore, the performance increases when features of the proposed (ST) and baseline (N-gram) are combined for the classification task. In addition, The proposed techniques have also been applied to the UPPC corpus to check their performance at the document level. The best result we obtained using the feature fusion technique ( F 1 = 0.855). Our corpus is available and free to download for research purposes.
Hamza Hafeez, Iqra Muneer, Muhammad Sharjeel, Muhammad Adnan Ashraf, Rao Muhammad Adeel Nawab
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2023 Developing a Large Benchmark Corpus for Urdu Semantic Word Similarity
abstract
The semantic word similarity task aims to quantify the degree of similarity between a pair of words. In literature, efforts have been made to create standard evaluation resources to develop, evaluate, and compare various methods for semantic word similarity. The majority of these efforts focused on English and some other languages. However, the problem of semantic word similarity has not been thoroughly explored for South Asian languages, particularly Urdu. To fill this gap, this study presents a large benchmark corpus of 518 word pairs for the Urdu semantic word similarity task, which were manually annotated by 12 annotators. To demonstrate how our proposed corpus can be used for the development and evaluation of Urdu semantic word similarity systems, we applied two state-of-the-art methods: (1) a word embedding–based method and (2) a Sentence Transformer–based method. As another major contribution, we proposed a feature fusion method based on Sentence Transformers and word embedding methods. The best results were obtained using our proposed feature fusion method (the combination of best features of both methods) with a Pearson correlation score of 0.67. To foster research in Urdu (an under-resourced language), our proposed corpus will be free and publicly available for research purposes.
Iqra Muneer, Ghazeefa Fatima, Muhammad Salman Khan 0001, Rao Muhammad Adeel Nawab, Ali Saeed
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 Semantic Tagging for the Urdu Language: Annotated Corpus and Multi-Target Classification Methods
abstract
Extracting and analysing meaning-related information from natural language data has attracted the attention of researchers in various fields, such as natural language processing, corpus linguistics, information retrieval, and data science. An important aspect of such automatic information extraction and analysis is the annotation of language data using semantic tagging tools. Different semantic tagging tools have been designed to carry out various levels of semantic analysis, for instance, named entity recognition and disambiguation, sentiment analysis, word sense disambiguation, content analysis, and semantic role labelling. Common to all of these tasks, in the supervised setting, is the requirement for a manually semantically annotated corpus, which acts as a knowledge base from which to train and test potential word and phrase-level sense annotations. Many benchmark corpora have been developed for various semantic tagging tasks, but most are for English and other European languages. There is a dearth of semantically annotated corpora for the Urdu language, which is widely spoken and used around the world. To fill this gap, this study presents a large benchmark corpus and methods for the semantic tagging task for the Urdu language. The proposed corpus contains 8,000 tokens in the following domains or genres: news, social media, Wikipedia, and historical text (each domain having 2K tokens). The corpus has been manually annotated with 21 major semantic fields and 232 sub-fields with the USAS (UCREL Semantic Analysis System) semantic taxonomy which provides a comprehensive set of semantic fields for coarse-grained annotation. Each word in our proposed corpus has been annotated with at least one and up to nine semantic field tags to provide a detailed semantic analysis of the language data, which allowed us to treat the problem of semantic tagging as a supervised multi-target classification task. To demonstrate how our proposed corpus can be used for the development and evaluation of Urdu semantic tagging methods, we extracted local, topical and semantic features from the proposed corpus and applied seven different supervised multi-target classifiers to them. Results show an accuracy of 94% on our proposed corpus which is free and publicly available to download.
Jawad Shafi, Rao Muhammad Adeel Nawab, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 Cross-lingual Text Reuse Detection at Document Level for English-Urdu Language Pair
abstract
In recent years, the problem of Cross-Lingual Text Reuse Detection (CLTRD) has gained the interest of the research community due to the availability of large digital repositories and automatic Machine Translation (MT) systems. These systems are readily available and openly accessible, which makes it easier to reuse text across languages but hard to detect. In previous studies, different corpora and methods have been developed for CLTRD at the sentence/passage level for the English-Urdu language pair. However, there is a lack of large standard corpora and methods for CLTRD for the English-Urdu language pair at the document level. To overcome this limitation, the significant contribution of this study is the development of a large benchmark cross-lingual (English-Urdu) text reuse corpus, called the TREU (Text Reuse for English-Urdu) corpus. It contains English to Urdu real cases of text reuse at the document level. The corpus is manually labelled into three categories (Wholly Derived = 672, Partially Derived = 888, and Non Derived = 697) with the source text in English and the derived text in the Urdu language. Another contribution of this study is the evaluation of the TREU corpus using a diversified range of methods to show its usefulness and how it can be utilized in the development of automatic methods for measuring cross-lingual (English-Urdu) text reuse at the document level. The best evaluation results, for both binary ( F 1 = 0.78) and ternary ( F 1 = 0.66) classification tasks, are obtained using a combination of all Translation plus Mono-lingual Analysis (T+MA) based methods. The TREU corpus is publicly available to promote CLTRD research in an under-resourced language, i.e., Urdu.
Muhammad Sharjeel, Iqra Muneer, Sumaira Nosheen, Rao Muhammad Adeel Nawab, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2022 Cross-Lingual Text Reuse Detection at sentence level for English-Urdu language pair
Iqra Muneer, Rao Muhammad Adeel Nawab
Comput. Speech Lang.2
2022 Developing a Cross-lingual Semantic Word Similarity Corpus for English-Urdu Language Pair
abstract
Semantic word similarity is a quantitative measure of how much two words are contextually similar. Evaluation of semantic word similarity models requires a benchmark corpus. However, despite the millions of speakers and the large digital text of the Urdu language on the Internet, there is a lack of benchmark corpus for the Cross-lingual Semantic Word Similarity task for the Urdu language. This article reports our efforts in developing such a corpus. The newly developed corpus is based on the SemEval-2017 task 2 English dataset, and it contains 1,945 cross-lingual English–Urdu word pairs. For each of these pairs of words, semantic similarity scores were assigned by 11 native Urdu speakers. In addition to corpus generation, this article also reports the evaluation results of a baseline approach, namely “Translation Plus Monolingual Analysis” for automated identification of semantic similarity between English–Urdu word pairs. The results showed that the path length similarity measure performs better for the Google and Bing translated words. The newly created corpus and evaluation results are freely available online for further research and development.
Ghazeefa Fatima, Rao Muhammad Adeel Nawab, Muhammad Salman Khan 0001, Ali Saeed
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Cross-lingual Text Reuse Detection Using Translation Plus Monolingual Analysis for English-Urdu Language Pair
abstract
Cross-Lingual Text Reuse Detection (CLTRD) has recently attracted the attention of the research community due to a large amount of digital text readily available for reuse in multiple languages through online digital repositories. In addition, efficient machine translation systems are freely and readily available to translate text from one language into another, which makes it quite easy to reuse text across languages, and consequently difficult to detect it. In the literature, the most prominent and widely used approach for CLTRD is Translation plus Monolingual Analysis (T+MA). To detect CLTR for English-Urdu language pair, T+MA has been used with lexical approaches, namely, N-gram Overlap, Longest Common Subsequence, and Greedy String Tiling. This clearly shows that T+MA has not been thoroughly explored for the English-Urdu language pair. To fulfill this gap, this study presents an in-depth and detailed comparison of 26 approaches that are based on T+MA. These approaches include semantic similarity approaches (semantic tagger based approaches, WordNet-based approaches), probabilistic approach (Kullback-Leibler distance approach), monolingual word embedding-based approaches siamese recurrent architecture, and monolingual sentence transformer-based approaches for English-Urdu language pair. The evaluation was carried out using the CLEU benchmark corpus, both for the binary and the ternary classification tasks. Our extensive experimentation shows that our proposed approach that is a combination of 26 approaches obtained an F 1 score of 0.77 and 0.61 for the binary and ternary classification tasks, respectively, and outperformed the previously reported approaches [ 41 ] ( F 1 = 0.73) for the binary and ( F 1 = 0.55) for the ternary classification tasks) on the CLEU corpus.
Iqra Muneer, Rao Muhammad Adeel Nawab
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Investigating the Feasibility of Deep Learning Methods for Urdu Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD), the process of automatically identifying the correct meaning of a word used in a given context, is a significant challenge in Natural Language Processing. A range of approaches to the problem has been explored by the research community. The majority of these efforts has focused on a relatively small set of languages, particularly English. Research on WSD for South Asian languages, particularly Urdu, is still in its infancy. In recent years, deep learning methods have proved to be extremely successful for a range of Natural Language Processing tasks. The main aim of this study is to apply, evaluate, and compare a range of deep learning methods approaches to Urdu WSD (both Lexical Sample and All-Words) including Simple Recurrent Neural Networks, Long-Short Term Memory, Gated Recurrent Units, Bidirectional Long-Short Term Memory, and Ensemble Learning. The evaluation was carried out on two benchmark corpora: (1) the ULS-WSD-18 corpus and (2) the UAW-WSD-18 corpus. Results (Accuracy = 63.25% and F1-Measure = 0.49) show that a deep learning approach outperforms previously reported results for the Urdu All-Words WSD task, whereas performance using deep learning approaches (Accuracy = 72.63% and F1-Measure = 0.60) are low in comparison to previously reported for the Urdu Lexical Sample task.
Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Sentiment analysis for Urdu online reviews using deep learning models
abstract
Abstract Most existing studies are focused on popular languages like English, Spanish, Chinese, Japanese, and others, however, limited attention has been paid to Urdu despite having more than 60 million native speakers. In this paper, we develop a deep learning model for the sentiments expressed in this under‐resourced language. We develop an open‐source corpus of 10,008 reviews from 566 online threads on the topics of sports, food, software, politics, and entertainment. The objectives of this work are bi‐fold (a) the creation of a human‐annotated corpus for the research of sentiment analysis in Urdu; and (b) measurement of up‐to‐date model performance using a corpus. For their assessment, we performed binary and ternary classification studies utilizing another model, namely long short‐term memory (LSTM), recurrent convolutional neural network (RCNN) Rule‐Based, N‐gram, support vector machine , convolutional neural network, and LSTM. The RCNN model surpasses standard models with 84.98% accuracy for binary classification and 68.56% accuracy for ternary classification. To facilitate other researchers working in the same domain, we have open‐sourced the corpus and code developed for this research.
Iqra Safder, Zainab Mahmood, Raheem Sarwar, Saeed-Ul Hassan, Farooq Zaman, Rao Muhammad Adeel Nawab, Faisal Bukhari, Rabeeh Ayaz Abbasi, Salem Alelyani, Naif R. Aljohani, Raheel Nawaz
Expert Syst. J. Knowl. Eng.6
2020 Deep sentiments in Roman Urdu text using Recurrent Convolutional Neural Network model
Zainab Mahmood, Iqra Safder, Rao Muhammad Adeel Nawab, Faisal Bukhari, Raheel Nawaz, Ahmed S. Alfakeeh, Naif R. Aljohani, Saeed-Ul Hassan
Inf. Process. Manag.3
2019 CLEU - A Cross-language english-urdu corpus and benchmark for text reuse experiments
abstract
Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large‐scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross‐language sentence/passage level text reuse corpus for the English‐Urdu language pair. The Cross‐Language English‐Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono‐lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross‐language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross‐language text reuse detection systems for the English‐Urdu language pair.
Iqra Muneer, Muhammad Sharjeel, Muntaha Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson
J. Assoc. Inf. Sci. Technol.4
2019 On comparing manual and automatic generated textual descriptions of business process models
abstract
Abstract Several organizations maintain textual process descriptions alongside graphical process descriptions to make them usable for all stakeholders. Maintaining textual process descriptions in the presence of continuously changing processes is a labor‐intensive task. Therefore, the automatic generation of textual descriptions is desirable. However, the trade‐offs between the manual and automatic generation of descriptions are yet to be investigated. To that end, this paper aims to answer two vital questions. How similar are the descriptions generated by the two approaches? What is the impact of using the two types of descriptions on process matching? To answer these specific questions, we have generated textual descriptions of 552 process models using the two approaches. To answer the first question, we have applied six text‐matching techniques and established that the descriptions overlap significantly; however, the formulation of sentences is substantially different. For answering the second question, we have used 11 text‐matching techniques to evaluate the impact of both descriptions on process matching. Results show (a) the choice of matching technique, and the type of description, have an impact on the matching performance and (b) vector space model (VSM) is the most appropriate matching technique whereas 5 gram is the worst performing technique.
Khurram Shahzad 0002, Sheeza Zaheer, Rao Muhammad Adeel Nawab, Faisal Aslam
J. Softw. Evol. Process.3
2019 A Sense Annotated Corpus for All-Words Urdu Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD) aims to automatically predict the correct sense of a word used in a given context. All human languages exhibit word sense ambiguity, and resolving this ambiguity can be difficult. Standard benchmark resources are required to develop, compare, and evaluate WSD techniques. These are available for many languages, but not for Urdu, despite this being a language with more than 300 million speakers and large volumes of text available digitally. To fill this gap, this study proposes a novel benchmark corpus for the Urdu All-Words WSD task. The corpus contains 5,042 words of Urdu running text in which all ambiguous words (856 instances) are manually tagged with senses from the Urdu Lughat dictionary. A range of baseline WSD models based on n -gram are applied to the corpus, and the best performance (accuracy of 57.71%) is achieved using word 4-gram. The corpus is freely available to the research community to encourage further WSD research in Urdu.
Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson 0001, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2017 Multilingual author profiling on Facebook
Mehwish Fatima, Komal Hasan, Saba Anwar, Rao Muhammad Adeel Nawab
Inf. Process. Manag.4
2017 An IR-Based Approach Utilizing Query Expansion for Plagiarism Detection in MEDLINE
abstract
The identification of duplicated and plagiarized passages of text has become an increasingly active area of research. In this paper, we investigate methods for plagiarism detection that aim to identify potential sources of plagiarism from MEDLINE, particularly when the original text has been modified through the replacement of words or phrases. A scalable approach based on Information Retrieval is used to perform candidate document selection-the identification of a subset of potential source documents given a suspicious text-from MEDLINE. Query expansion is performed using the ULMS Metathesaurus to deal with situations in which original documents are obfuscated. Various approaches to Word Sense Disambiguation are investigated to deal with cases where there are multiple Concept Unique Identifiers (CUIs) for a given term. Results using the proposed IR-based approach outperform a state-of-the-art baseline based on Kullback-Leibler Distance.
Rao Muhammad Adeel Nawab, Mark Stevenson 0001, Paul D. Clough
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Urdu Summary Corpus
Muhammad Humayoun, Rao Muhammad Adeel Nawab, Saba Aslam, Omer Farzand
LREC2
2016 Lexical Coverage Evaluation of Large-scale Multilingual Semantic Lexicons for Twelve Languages
Scott Piao, Paul Rayson, Dawn Archer, Francesca Bianchi, Carmen Dayrell, Mahmoud El-Haj, Ricardo-María Jiménez, Dawn Knight, Michal Kren, Laura Löfberg, Rao Muhammad Adeel Nawab, Jawad Shafi, Phoey Lee Teh, Olga Mudraya
LREC11
2016 UPPC - Urdu Paraphrase Plagiarism Corpus
Muhammad Sharjeel, Paul Rayson, Rao Muhammad Adeel Nawab
LREC3
2014 Comparing Medline citations using modified N-grams
abstract
OBJECTIVE: We aim to identify duplicate pairs of Medline citations, particularly when the documents are not identical but contain similar information. MATERIALS AND METHODS: Duplicate pairs of citations are identified by comparing word n-grams in pairs of documents. N-grams are modified using two approaches which take account of the fact that the document may have been altered. These are: (1) deletion, an item in the n-gram is removed; and (2) substitution, an item in the n-gram is substituted with a similar term obtained from the Unified Medical Language System Metathesaurus. N-grams are also weighted using a score derived from a language model. Evaluation is carried out using a set of 520 Medline citation pairs, including a set of 260 manually verified duplicate pairs obtained from the Deja Vu database. RESULTS: The approach accurately detects duplicate Medline document pairs with an F1 measure score of 0.99. Allowing for word deletions and substitution improves performance. The best results are obtained by combining scores for n-grams of length 1-5 words. DISCUSSION: Results show that the detection of duplicate Medline citations can be improved by modifying n-grams and that high performance can also be obtained using only unigrams (F1=0.959), particularly when allowing for substitutions of alternative phrases.
Rao Muhammad Adeel Nawab, Mark Stevenson 0001, Paul D. Clough
J. Am. Medical Informatics Assoc.1
2012 Retrieving Candidate Plagiarised Documents Using Query Expansion
Rao Muhammad Adeel Nawab, Mark Stevenson 0001, Paul D. Clough
ECIR1