Saeed-Ul Hassan

dblp:78/11279 · DBLP profile ↗
← Back
29ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0002-6509-9190ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
citation analysis
0.412020
Automatic Classification of Algorithm Citation Functions in Scientific Literature · IEEE Trans. Knowl. Data Eng. 2020
Empirical software engineering
mining software repositories
0.412020
Automatic Classification of Algorithm Citation Functions in Scientific Literature · IEEE Trans. Knowl. Data Eng. 2020

Methods — techniques the papers use, named apart from their topics

heterogeneous ensemble machine learning · 0.9base classifiers · 0.9
YearPublicationVenuePosition
2025 Urdu paraphrased text reuse and plagiarism detection using pre-trained large language models and deep hybrid neural networks
abstract
The growing prevalence of text reuse and plagiarism in various fields has led to an urgent need for reliable computational methods for detection. However, current commercial plagiarism detection systems are ineffective in identifying paraphrased cases of text reuse, highlighting the need for improvement. Previous research on paraphrased text reuse and plagiarism detection has mainly focused on English, European, Persian, and Arabic languages, and very few studies have been reported on the under-resourced Urdu language. This study aims to overcome this research gap by using a Deep Neural Network (DNN) based architecture and pre-trained Large Language Models (LLMs) for the task of Urdu paraphrased text reuse and plagiarism detection. The architecture called Deep Text Reuse and Paraphrased Plagiarism Detection (D-TRaPPD), relies on LLMs for input and utilizes CNN and LSTM to extract essential textual features. Moreover, we have proposed and evaluated two D-TRaPPD variants, Word Embeddings-D-TRaPPD (WE-D-TRaPPD) and Sentence Embeddings-D-TRaPPD (SE-D-TRaPPD), using two gold standard document-level corpora containing both real and simulated cases of Urdu paraphrased text reuse and plagiarism. The results demonstrate the effectiveness of the D-TRaPPD architecture, with SE-D-TRaPPD achieving the highest $$F_1$$ scores of 91.77 for real cases and 95.15 for simulated cases. Furthermore, the results highlight the superiority of our approaches over the state-of-the-art methods for Urdu paraphrased text reuse and plagiarism detection.
Hafiz Rizwan Iqbal, Muhammad Sharjeel, Jawad Shafi, Usama Mehmood, Saeed-Ul Hassan, Agha Ali Raza
Multim. Tools Appl.5
2025 UPON: Urdu Poetry Generation Using Deep Learning: A Novel Approach and Evaluation
abstract
Poetry represents the oldest and most esteemed literary form, allowing poets to convey ideas while carefully attending to elements such as meaning, coherence, poetic quality, and fluency. Notably, the creation of good poetry entails considerations of rhyme and meter. With the advent of artificial intelligence (AI), significant advancements have been made in automatic text generation, primarily within languages such as English and Chinese. However, the generation of Urdu poetry presents a unique challenge due to the language’s inherent ambiguity, cultural and historical nuances, and the demand for creativity. The existing body of literature has only marginally explored Urdu prose and has almost entirely overlooked the domain of Urdu poetry generation, primarily due to the scarcity of comprehensive training data. In response to this deficiency, this research endeavor addresses this challenge. It begins by introducing a specialized Urdu poetry dataset adhering to a specific meter, “behr-e-khafeef,” which incorporates approximately 20,000 couplets from the Rekhta repository. Subsequently, a character-based encoding methodology is proposed to transform these couplets into a numerical representation, assigning a distinct identifier to each character. The generation process initiates with the creation of the first verse through a character-level LSTM, followed by the application of a machine translation technique, specifically sequence-to-sequence learning, to formulate the second verse based on the first. The generated poetry is subjected to evaluation based on metrics, including BLEU scores. Additionally, an expert panel of Urdu poets is engaged to conduct a human assessment of the generated couplets, with the evaluation encompassing critical dimensions such as meaning, coherence, poetic quality, and fluency. Our findings are juxtaposed with existing poetry generation systems, demonstrating a notable advancement in the state-of-the-art, as evidenced by a BLEU score of 0.23. The research culminates with the presentation of prospective avenues for further exploration, aimed at inspiring the scholarly community to enhance the domain of poetry generation and augment existing contributions in this field.
Muhammad Rauf Tabassam, Hajra Waheed, Iqra Safder, Raheem Sarwar, Naif R. Aljohani, Raheel Nawaz, Saeed-Ul Hassan, Farooq Zaman, Muhammad Ahtazaz Ahsan
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2024 Urdu paraphrase detection: A novel DNN-based implementation using a semi-automatically generated corpus
abstract
Abstract Automatic paraphrase detection is the task of measuring the semantic overlap between two given texts. A major hurdle in the development and evaluation of paraphrase detection approaches, particularly for South Asian languages like Urdu, is the inadequacy of standard evaluation resources. The very few available paraphrased corpora for these languages are manually created. As a result, they are constrained to smaller sizes and are not very feasible to evaluate mainstream data-driven and deep neural networks (DNNs)-based approaches. Consequently, there is a need to develop semi- or fully automated corpus generation approaches for the resource-scarce languages. There is currently no semi- or fully automatically generated sentence-level Urdu paraphrase corpus. Moreover, no study is available to localize and compare approaches for Urdu paraphrase detection that focus on various mainstream deep neural architectures and pretrained language models. This research study addresses this problem by presenting a semi-automatic pipeline for generating paraphrased corpora for Urdu. It also presents a corpus that is generated using the proposed approach. This corpus contains 3147 semi-automatically extracted Urdu sentence pairs that are manually tagged as paraphrased (854) and non-paraphrased (2293). Finally, this paper proposes two novel approaches based on DNNs for the task of paraphrase detection in Urdu text. These are Word Embeddings n-gram Overlap (henceforth called WENGO), and a modified approach, Deep Text Reuse and Paraphrase Plagiarism Detection (henceforth called D-TRAPPD). Both of these approaches have been evaluated on two related tasks: (i) paraphrase detection, and (ii) text reuse and plagiarism detection. The results from these evaluations revealed that D-TRAPPD ( $F_1 = 96.80$ for paraphrase detection and $F_1 = 88.90$ for text reuse and plagiarism detection) outperformed WENGO ( $F_1 = 81.64$ for paraphrase detection and $F_1 = 61.19$ for text reuse and plagiarism detection) as well as other state-of-the-art approaches for these two tasks. The corpus, models, and our implementations have been made available as free to download for the research community.
Hafiz Rizwan Iqbal, Rashad Maqsood, Agha Ali Raza, Saeed-Ul Hassan
Nat. Lang. Eng.4
2024 MSDGSD: A Scalable Graph Descriptor for Processing Large Graphs
abstract
Graph representation methods have recently become the de facto standard for downstream machine learning tasks on graph-structured data and have found numerous applications, e.g., drug discovery & development, recommendation, and forecasting. However, the existing methods are specially designed to work in a centralized environment, which limits their applicability to small or medium-sized graphs. In this work, we present a graph embedding method that extracts graph representations in a distributed environment with independent and parallel machines. The proposed method is built-upon the existing approach, distributed graph statistical distance (DGSD), to enhance the scalability on large graphs. The key innovation of our work lies in the proposition of a batching mechanism for client-server message passing, which reduces communication overhead during the computation of the distance matrix. In addition, we present a sampling approach for computing pairwise distances between the nodes to compute the desired graph embedding. Moreover, we systematically explore six distinct variations of a distributed graph embeddings and subsequently subject them to comprehensive evaluation. Our extensive evaluations on over 20 graph datasets and ten baseline methods demonstrate improved running time and comparative classification accuracy compared to state-of-the-art embedding techniques.
Anwar Said, Iqra Safder, Saeed-Ul Hassan, Naif R. Aljohani, Mudassir Shabbir
IEEE Trans. Comput. Soc. Syst.4
2023 Early prediction of learners at risk in self-paced education: A neural network approach
Hajra Waheed, Saeed-Ul Hassan, Raheel Nawaz, Naif R. Aljohani, Guanliang Chen, Dragan Gasevic
Expert Syst. Appl.2
2023 Neural machine translation for in-text citation classification
abstract
Abstract The quality of scientific publications can be measured by quantitative indices such as the h‐index, Source Normalized Impact per Paper, or g‐index. However, these measures lack to explain the function or reasons for citations and the context of citations from citing publication to cited publication. We argue that citation context may be considered while calculating the impact of research work. However, mining citation context from unstructured full‐text publications is a challenging task. In this paper, we compiled a data set comprising 9,518 citations context. We developed a deep learning‐based architecture for citation context classification. Unlike feature‐based state‐of‐the‐art models, our proposed focal‐loss and class‐weight‐aware BiLSTM model with pretrained GloVe embedding vectors use citation context as input to outperform them in multiclass citation context classification tasks. Our model improves on the baseline state‐of‐the‐art by achieving an F1 score of 0.80 with an accuracy of 0.81 for citation context classification. Moreover, we delve into the effects of using different word embeddings on the performance of the classification model and draw a comparison between fastText, GloVe, and spaCy pretrained word embeddings.
Iqra Safder, Momin Ali, Naif R. Aljohani, Raheel Nawaz, Saeed-Ul Hassan
J. Assoc. Inf. Sci. Technol.5
2023 Predicting functional roles of Ethereum blockchain addresses
Tania Saleem, Muhammad Ismaeel, Muhammad Umar Janjua, Abdulrahman Ali, Awab Aqib, Saeed-Ul Hassan
Peer Peer Netw. Appl.7
2022 Uncovering Associations Between Cognitive Presence and Speech Acts: A Network-Based Approach
abstract
This research aimed to explore the relationship between different indicators of the depth and quality of participation in computer-mediated learning environments. By using network analyses and statistical tests, we discovered significant associations between the cognitive presence phases of the Community of Inquiry framework and speech acts, and examined the impact of two different instructional interventions on these associations. We found that there are strong associations between some speech acts and cognitive presence phases. In addition, the study revealed that the association between speech acts and cognitive presence is moderated by external facilitation, but not affected by user role assignment. The results suggest that speech acts can plausibly be used to provide feedback in relation to cognitive presence and can potentially be used to increase the generalizability of cognitive presence classification.
Sehrish Iqbal, Zach Swiecki, Srecko Joksimovic, Rafael Ferreira Leite de Mello, Naif R. Aljohani, Saeed-Ul Hassan, Dragan Gasevic
LAK6
2022 Monitoring the Growth Status of Corn Crop from UAV Images Based on Dense Convolutional Neural Network
abstract
Monitoring corn crop growth status is of great significance to crop production, breeding, and seed production. The Unmanned Aerial Vehicles’ (UAVs) technology makes it possible to use computer vision technology to identify corn growth stage intelligently. A model customized for corn growth status monitoring based on a dense convolutional neural network (CM-CNN) was proposed, including a two-way dense module and a new activation function ELU. The two-way dense module enlarges the receptive field, while the ELU alleviates gradient disappearance and speeds up learning in deep neural networks. Dense architecture concatenates all the previous layer features to enhance feature reuse. The proposed CM-CNN performs well in classifying corn growth stages. Experimental results show that CM-CNN is a state-of-the-art method, with an accuracy of its relevant data up to 99.3%. Compared with other CNN models, viz. AlexNet, ZFNet, VGG, InceptionV3, Xception and ResNet, fewer parameters are in CM-CNN.
Jia Zhu 0003, Yuling Xing, Zhangyan Dai, Jin Huang 0007, Saeed-Ul Hassan
Int. J. Pattern Recognit. Artif. Intell.6
2022 Predictive Model Using a Machine Learning Approach for Enhancing the Retention Rate of Students At-Risk
abstract
Student retention is a widely recognized challenge in the educational community to assist the institutes in the formation of appropriate and effective pedagogical interventions. This study intends to predict the students at-risk of low performances during an on-going course, those at-risk of graduating late than the tentative timeline and predicting the capacity of students in a campus. The data constitutes of demographics, learning, academic and educational related attributes which are suitable to deploy various machine learning algorithms for the prediction of at-risk students. For class balancing, Synthetic Minority Over Sampling Technique, is also applied to eliminate the imbalance in the academic award-gap performances and late/timely graduates. Results reveal the effectiveness of the deployed techniques with Long short-term Memory (LSTM) outperforming other models for early prediction of at-risk students. The main contribution of this work is a machine learning approach capable of enhancing the academic decision making related to student performance.
Hani Brdesee, Wafaa Adnan Alsaggaf, Naif R. Aljohani, Saeed-Ul Hassan
Int. J. Semantic Web Inf. Syst.4
2022 A dataset and benchmark for malaria life-cycle classification in thin blood smear images
Qazi Ammar Arshad, Mohsen Ali, Saeed-Ul Hassan, Chen Chen 0001, Ayisha Imran, Ghulam Rasul, Waqas Sultani
Neural Comput. Appl.3
2022 UrduAI: Writeprints for Urdu Authorship Identification
abstract
The authorship identification task aims at identifying the original author of an anonymous text sample from a set of candidate authors. It has several application domains such as digital text forensics and information retrieval. These application domains are not limited to a specific language. However, most of the authorship identification studies are focused on English and limited attention has been paid to Urdu. However, existing Urdu authorship identification solutions drop accuracy as the number of training samples per candidate author reduces and when the number of candidate authors increases. Consequently, these solutions are inapplicable to real-world cases. Moreover, due to the unavailability of reliable POS taggers or sentence segmenters, all existing authorship identification studies on Urdu text are limited to the word n-grams features only. To overcome these limitations, we formulate a stylometric feature space, which is not limited to the word n-grams feature only. Based on this feature space, we use an authorship identification solution that transforms each text sample into a point set, retrieves candidate text samples, and relies on the nearest neighbors classifier to predict the original author of the anonymous text sample. To evaluate our solution, we create a significantly larger corpus than existing studies and conduct several experimental studies that show that our solution can overcome the limitations of existing studies and report an accuracy level of 94.03%, which is higher than all previous authorship identification works.
Raheem Sarwar, Saeed-Ul Hassan
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Sentiment analysis for Urdu online reviews using deep learning models
abstract
Abstract Most existing studies are focused on popular languages like English, Spanish, Chinese, Japanese, and others, however, limited attention has been paid to Urdu despite having more than 60 million native speakers. In this paper, we develop a deep learning model for the sentiments expressed in this under‐resourced language. We develop an open‐source corpus of 10,008 reviews from 566 online threads on the topics of sports, food, software, politics, and entertainment. The objectives of this work are bi‐fold (a) the creation of a human‐annotated corpus for the research of sentiment analysis in Urdu; and (b) measurement of up‐to‐date model performance using a corpus. For their assessment, we performed binary and ternary classification studies utilizing another model, namely long short‐term memory (LSTM), recurrent convolutional neural network (RCNN) Rule‐Based, N‐gram, support vector machine , convolutional neural network, and LSTM. The RCNN model surpasses standard models with 84.98% accuracy for binary classification and 68.56% accuracy for ternary classification. To facilitate other researchers working in the same domain, we have open‐sourced the corpus and code developed for this research.
Iqra Safder, Zainab Mahmood, Raheem Sarwar, Saeed-Ul Hassan, Farooq Zaman, Rao Muhammad Adeel Nawab, Faisal Bukhari, Rabeeh Ayaz Abbasi, Salem Alelyani, Naif R. Aljohani, Raheel Nawaz
Expert Syst. J. Knowl. Eng.4
2021 Automatic team recommendation for collaborative software development
Suppawong Tuarob, Noppadol Assavakamhaenghan, Waralee Tanaphantaruk, Ponlakit Suwanworaboon, Saeed-Ul Hassan, Morakot Choetkiertikul
Empir. Softw. Eng.5
2021 DGSD: Distributed graph representation via graph statistical properties
Anwar Said, Saeed-Ul Hassan, Suppawong Tuarob, Raheel Nawaz, Mudassir Shabbir
Future Gener. Comput. Syst.2
2021 NetKI: A kirchhoff index based statistical graph embedding in nearly linear time
Anwar Said, Saeed-Ul Hassan, Waseem Abbas 0003, Mudassir Shabbir
Neurocomputing2
2020 Deep sentiments in Roman Urdu text using Recurrent Convolutional Neural Network model
Zainab Mahmood, Iqra Safder, Rao Muhammad Adeel Nawab, Faisal Bukhari, Raheel Nawaz, Ahmed S. Alfakeeh, Naif R. Aljohani, Saeed-Ul Hassan
Inf. Process. Manag.8
2020 Deep Learning-based Extraction of Algorithmic Metadata in Full-Text Scholarly Documents
Iqra Safder, Saeed-Ul Hassan, Anna Visvizi, Thanapon Noraset, Raheel Nawaz, Suppawong Tuarob
Inf. Process. Manag.2
2020 HTSS: A novel hybrid text summarisation and simplification architecture
Farooq Zaman, Matthew Shardlow, Saeed-Ul Hassan, Naif R. Aljohani, Raheel Nawaz
Inf. Process. Manag.3
2020 Predicting literature's early impact with sentiment analysis in Twitter
Saeed-Ul Hassan, Naif R. Aljohani, Nimra Idrees, Raheem Sarwar, Raheel Nawaz, Eugenio Martínez-Cámara, Sebastián Ventura, Francisco Herrera
Knowl. Based Syst.1
2020 Bot prediction on social networks of Twitter in altmetrics using deep graph convolutional networks
Naif R. Aljohani, Ayman G. Fayoumi, Saeed-Ul Hassan
Soft Comput.3
2020 ArWordVec: efficient word embedding models for Arabic tweets
Mohammed M. Fouad 0002, Ahmed Mahany, Naif R. Aljohani, Rabeeh Ayaz Abbasi, Saeed-Ul Hassan
Soft Comput.5
2020 Native Language Identification of Fluent and Advanced Non-Native Writers
abstract
Native Language Identification (NLI) aims at identifying the native languages of authors by analyzing their text samples written in a non-native language. Most existing studies investigate this task for educational applications such as second language acquisition and require the learner corpora. This article performs NLI in a challenging context of the user-generated-content (UGC) where authors are fluent and advanced non-native speakers of a second language. Existing NLI studies with UGC (i) rely on the content-specific/social-network features and may not be generalizable to other domains and datasets, (ii) are unable to capture the variations of the language-usage-patterns within a text sample, and (iii) are not associated with any outlier handling mechanism. Moreover, since there is a sizable number of people who have acquired non-English second languages due to the economic and immigration policies, there is a need to gauge the applicability of NLI with UGC to other languages. Unlike existing solutions, we define a topic-independent feature space, which makes our solution generalizable to other domains and datasets. Based on our feature space, we present a solution that mitigates the effect of outliers in the data and helps capture the variations of the language-usage-patterns within a text sample. Specifically, we represent each text sample as a point set and identify the top- k stylistically similar text samples (SSTs) from the corpus. We then apply the probabilistic k nearest neighbors’ classifier on the identified top- k SSTs to predict the native languages of the authors. To conduct experiments, we create three new corpora where each corpus is written in a different language, namely, English, French , and German . Our experimental studies show that our solution outperforms competitive methods and reports more than 80% accuracy across languages.
Raheem Sarwar, Attapol Rutherford, Saeed-Ul Hassan, Thanawin Rakthanmanon, Sarana Nutanong
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 Automatic Classification of Algorithm Citation Functions in Scientific Literature
abstract
Computer sciences and related disciplines evolve around developing, evaluating, and applying algorithms. Typically, an algorithm is not developed from scratch, but uses and builds upon existing ones, which often are proposed and published in scholarly articles. The ability to capture this evolution relationship among these algorithms in scientific literature would not only allow us to understand how a particular algorithm is composed, but also shed light on large-scale analysis of algorithmic evolution through different temporal spans and thematic scales. We propose to capture such evolution relationship between two algorithms by investigating the knowledge represented in citation contexts, where authors explain how cited algorithms are used in their works. A set of heterogeneous ensemble machine-learning methods is proposed, where the combination of two base classifiers trained with heterogeneous feature types is used to automatically identify the algorithm usage relationship. The proposed heterogeneous ensemble methods achieve the best average F1 of 0.749 and 0.905 for fine-grained and binary algorithm citation function classification, respectively. The success of this study will allow us to generate a large-scale algorithm citation network from a collection of scholarly documents representing multiple time spans, venues, and fields of study. Such a network will be used as an instrument not only to answer critical questions in algorithm search, such as identifying the most influential and generalizable algorithms, but also to study the evolution of algorithmic development and trends over time.
Suppawong Tuarob, Sung Woo Kang, Poom Wettayakorn, Chanathip Pornprasit, Tanakitti Sachati, Saeed-Ul Hassan, Peter Haddawy
IEEE Trans. Knowl. Data Eng.6
2019 Urdu language based information dissemination system for low-literate farmers
abstract
This paper describes the design process by which we designed an Android application equipped with audio, textual menus and visuals components for use by farmers of diverse literacy levels looking for vital weather information after the conclusion of research-work that productivity lags due to information inadequacies. The intervention provides more timely access to accurate information to low-literate farmers and thereby help in making the agricultural ecosystem more robust. We discuss the various design and implementation features of our system and presents our findings from the field on the usability of our application. We have also openly released our source code so that other users and developers can also benefit from our work.
Fahad Idrees, Junaid Qadir 0001, Hamid Mehmood, Saeed-Ul Hassan, Amna Batool
ICTD4
2019 The 'who' and the 'what' in international migration research: data-driven analysis of Scopus-indexed scientific literature
abstract
This paper offers a detailed, first-ever, in-depth, data-driven, review of debates pertaining to international migration (IM) as depicted by cross-disciplinary records collected in Scopus. Accordingly, the paper also makes a case for the value added of bibliometric analysis and new ways of its application. Specifically, to gain a thorough understanding of issues, names, and topics that have contributed to the IM debate since 1963, bibliometric analysis was conducted on 12.663 procured records. The findings suggest that regardless of the depth and breadth of the analysis, it is doomed to remain partial. That is, when confronted with academic work not available in Scopus, this study concludes that more work needs to be done to ensure, on the one hand, interoperability of research data repositories, and on the other hand, synergies among the until now divided research-communities research and publishing in the field of IM. Only in this way, it is argued, will it be possible to ensure transparency of research artefacts, identify issues and problems silenced and/or under-researched in the field, and finally enable more efficient dialogue between academia and decision-makers, including international organisations.
Saeed-Ul Hassan, Anna Visvizi, Hajra Waheed
Behav. Inf. Technol.1
2019 Virtual learning environment to predict withdrawal by leveraging deep learning
abstract
The current evolution in multidisciplinary learning analytics research poses significant challenges for the exploitation of behavior analysis by fusing data streams toward advanced decision-making. The identification of students that are at risk of withdrawals in higher education is connected to numerous educational policies, to enhance their competencies and skills through timely interventions by academia. Predicting student performance is a vital decision-making problem including data from various environment modules that can be fused into a homogenous vector to ascertain decision-making. This research study exploits a temporal sequential classification problem to predict early withdrawal of students, by tapping the power of actionable smart data in the form of students' interactional activities with the online educational system, using the freely available Open University Learning Analytics data set by employing deep long short-term memory (LSTM) model. The deployed LSTM model outperforms baseline logistic regression and artificial neural networks by 10.31% and 6.48% respectively with 97.25% learning accuracy, 92.79% precision, and 85.92% recall.
Saeed-Ul Hassan, Hajra Waheed, Naif R. Aljohani, Mohsen Ali, Sebastián Ventura, Francisco Herrera
Int. J. Intell. Syst.1
2019 Web Observatory Insights: Past, Current, and Future
abstract
In the present era of Big Data, with continuously increasing amounts of user-generated content, it is becoming a challenge to understand the relation between the content that is available on the Web and the users who are generating that content. Researchers have come up with many ways to understand today's Web better. One of the recently introduced concepts is a Web observatory (WO). This article provides a deep understanding about web observatories. It discusses the status of existing WO systems. The article investigates and gathers the common practices of WOs. This research has implications for researchers and communities in the adoption of the WO concept. The article highlights the challenges of WOs, such as data crawling, privacy and security. It also provides future research and development directions. The article provides a comparative analysis of existing WOs. It discusses the architecture of WOs. It presents components of a WO in a coherent manner and finally provides insights into challenges and limitations of WOs.
Naif R. Aljohani, Rabeeh Ayaz Abbasi, Fahad Mohammed Bawakid, Farrukh Saleem, Zahid Ullah 0004, Ali Daud, Muhammad Ahtisham Aslam, Jalal S. Alowibdi, Saeed-Ul Hassan
Int. J. Semantic Web Inf. Syst.9
2018 A bibliometric perspective of learning analytics research landscape
abstract
Learning analytics is an emerging field of research, motivated by the wide spectrum of the available educational information that can be analysed to provide a data-driven decision about various learning problems. This study intends to examine the research landscape of learning analytics to deliver a comprehensive understanding of the research activities in this multidisciplinary field, using scientific literature from the Scopus database. An array of state-of-the-art bibliometric indices is deployed on 2811 procured publication datasets: publication counts, citation counts, co-authorship patterns, citation networks and term co-occurrence. The results indicate that the field of learning analytics appears to have been instantiated around 2011; thus, before this time period no significant research activity can be observed. The temporal evolution indicates that the terms ‘students’, ‘teachers’, ‘higher education institutions’ and ‘learning process’ appear to be the major components of the field. More recent trends in the field are the tools that tap into Big Data analytics and data mining techniques for more rational data-driven decision-making services. A future direction research depicts a need to integrate learning analytics research with multidisciplinary smart education and smart library services. The vision towards smart city research requires a meta-level of smart learning analytics value integration and policy-making.
Hajra Waheed, Saeed-Ul Hassan, Naif R. Aljohani
Behav. Inf. Technol.2