Farhan Ullah 0001

dblp:63/171-1 · DBLP profile ↗
← Back
33ranked-venue papers
11as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TIDE-Net: A Two-Stage Temporal Deep Learning Framework for Multi-Granular IoT Intrusion Detection
abstract
ABSTRACT The rapid growth of the internet of things (IoT) has increased exposure to advanced cyberattacks. However, most existing intrusion detection systems (IDS) rely on outdated or synthetic datasets that do not reflect real deployment conditions. The recently released IDSIoT2024 dataset provides long‐term traffic traces from real IoT devices, allowing a more realistic evaluation of intrusion detection models. In this paper, we propose TIDE‐Net, a two‐stage temporal deep learning framework designed for the characteristics of IDSIoT2024. In Stage 1, the framework performs binary classification to separate benign and malicious traffic, whereas Stage 2 deals with the malicious traffic which is further classified into either three coarse‐grained attack categories or twelve fine‐grained attack types. Deep neural networks (DNN), one‐dimensional convolutional neural networks (CNN), and bidirectional long short‐term memory networks (BiLSTM) are evaluated under three settings: Binary, 3‐class, and 12‐class classification. Among these models, BiLSTM shows the most stable performance across all tasks. It achieves 99.24% accuracy in binary detection, over 98.7% accuracy in 3‐class classification, and a macro F1‐score of 0.914 in 12‐class classification. The proposed two‐stage BiLSTM‐based pipeline achieves approximately 97% end‐to‐end accuracy. It also handles class imbalance and temporal patterns more effectively than CNN and DNN baselines. These results provide one of the first comprehensive deep learning benchmarks on IDSIoT2024 and confirm the effectiveness of hierarchical, BiLSTM‐based temporal models for IoT intrusion detection.
Mirza Qais Baig, Ali Turab, Farhan Ullah 0001
Concurr. Comput. Pract. Exp.3
2026 PAACGAN : Protocol-Aware Generative Augmentation for Robust Imbalanced Network Intrusion Detection
abstract
ABSTRACT Class imbalance remains a major challenge in network intrusion detection because rare attack classes are substantially underrepresented relative to benign traffic and frequent attack categories. This imbalance biases learning algorithms towards majority classes and reduces detection performance on critical minority attacks. To address this problem, we propose PAACGAN, a Protocol‐Aware Auxiliary Classifier Generative Adversarial Network for minority‐class augmentation in intrusion detection. The proposed model incorporates protocol‐aware conditioning based on TCP/UDP characteristics, destination port distributions, and class‐specific traffic statistics so that generated samples preserve both class identity and protocol‐level consistency. The augmented training data are then used to train an XGBoost classifier for intrusion detection. Experiments on the NSL‐KDD and CICIDS2017 benchmark datasets show that the proposed framework improves recall and F1‐score for minority classes. On the rarest CICIDS2017 attack categories, including Infiltration and Heartbleed , the proposed method achieves an F1‐score improvement of up to 12.57 percentage points. These results show that protocol‐aware generative augmentation is effective for improving rare‐attack detection under severe data imbalance.
Mohyminul Islam, Ali Turab, Farhan Ullah 0001
Expert Syst. J. Knowl. Eng.3
2026 HoloTiny-AD: A Trustworthy Anomaly Detection in Resource-Constrained IoT Devices Using Holographic TinyML and Deep Metaheuristics
abstract
Mobile Edge Computing (MEC) security and performance may be threatened by network traffic anomalies such as illegal data transfers or access. Traditional detection methods, such as machine learning, statistical models, and signature-based methods, are ineffective for real-time monitoring in resource-constrained devices. TinyML deploys lightweight models on resource-limited devices for rapid, localized, low-power threat detection without cloud access, improving MEC real-time security. This paper introduces HoloTiny-AD, a real-time anomaly detection technique in MEC using the Holographic Counterpart (HC) architecture, TinyML, and Deep Particle Swarm Optimization (DPSO). The proposed approach effectively monitors and detects anomalies while using minimum computing resources, making it suitable for MEC. HC uses TinyML models trained with lightweight classifiers to mimic device state transitions and analyze anomalies to identify potential risks. DPSO optimizes parameters and task scheduling to increase detection accuracy and predict overload or busy state device failures. The proposed methodology is evaluated on two benchmark datasets: CIC-BCCC-NRC-ACI-IOT-2023 and CIC-IDS2017. Results show that HoloTiny-AD is a robust, scalable solution for securing edge devices against evolving network threats under resource constraints.
Yue Zhao 0014, Gautam Srivastava 0001, Farhan Ullah 0001
IEEE Internet Things J.3
2026 Leveraging Multimodal LLMs and Metaverse Technologies for Early Diagnosis of Elderly Diseases
Ahmad Nawaz Zaheer, Muhammad Jamil, Farhan Ullah 0001, Awais Ahmad 0001, Fakhri Alam Khan, Gwanggil Jeon, Sheeraz Akram
IEEE Trans. Comput. Soc. Syst.4
2026 MSEF-YOLO11s: a multi-scale extraction and fusion network for small target detection in drone imagery
abstract
Abstract Small object detection in unmanned aerial vehicle (UAV) aerial imagery faces substantial challenges due to small target scales, complex backgrounds, noise interference, and so on. To enhance multi-scale feature representation and detection efficiency, this paper proposes MSEF-YOLO11s. Specifically, we first design a lightweight partial multi-scale (LPMS) module, which effectively aggregates cross-scale information and enhances multi-scale representations in the backbone for small objects. Secondly, to dynamically adjust feature weights and mitigate feature conflicts in the neck, we devise a multi-scale boundary-semantic alignment (MS-BSA) based on adaptive attention, which can further avoid computational redundancy for sufficient fusion. Finally, a lightweight shared detail detection head (LSDDH) replaces the decoupled head structure with shared convolutional layers, resolving the issue of parameter explosion associated with adding a dedicated small object detection head. Experimental results demonstrate the effectiveness of the proposed model. Specifically, compared to the baseline YOLO11s, MSEF-YOLO11s achieves an improvement of 6.6% in mAP50 on the VisDrone2019 test set, with only 4.4M increase in parameters. Furthermore, mAP50 on the TinyPerson test set increases from 22.8% to 28.1%, confirming the model’s strong generalization capability.
Farhan Ullah 0001, Yue Zhao 0014
J. Supercomput.3
2025 A Deep Learning-Based RIDNet Approach for Enhanced Denoising of SAR Images
Muhammad Pervez Akhter, Farhan Ullah 0001, Leonardo Mostarda, Diletta Cacciagrano
AINA (7)2
2025 H_SIG: Privacy-Preserving Auction for Big Data Based on Homomorphic Signcryption
Shamsher Ullah, Farhan Ullah 0001, Muhammad Umar Farooq 0002, Gautam Srivastava 0001, Victor C. M. Leung
IEEE Big Data2
2025 Classification of intrusion cyber-attacks in smart power grids using deep ensemble learning with metaheuristic-based optimization
abstract
Abstract The most advanced power grid design, known as a ‘smart power grid’, integrates information and communication technology (ICT) with a conventional grid system to enable remote management of electricity distribution. The intelligent cyber‐physical architecture enables bidirectional, real‐time data sharing between electricity suppliers and consumers through smart meters and advanced metering infrastructure (AMI). Data protection issues, such as data tampering, firmware exploitation, and the leakage of sensitive information arise due to the smart power grid's substantial reliance on ICT. To maintain reliable and efficient power distribution, these issues must be identified and resolved quickly. Intrusion detection is essential for providing secure services and alerting system administrators in the case of adversary attacks. This paper proposes an intrusion classification scheme that identifies several types of cyber attacks on modern smart power grids. Grey‐Wolf metaheuristic optimization‐based feature selection is used to learn non‐linear, overlapping, and complex electrical grid properties. An extended deep‐stacked ensemble technique is advanced by putting predictions from weak learners (CNNs) into a meta‐learner (MLP). The outcomes of this approach are explained and confirmed using explainable AI (XAI). The publicly available dataset from Mississippi State University and Oak Ridge National Laboratory (MSU‐ORNL) is used to conduct experiments. The experimental results show that the proposed method achieved a peak accuracy of 96.6% while scrutinizing the original MSU‐ORNL data feature set and a maximum accuracy of 99% when analysing the selected feature set. Therefore, the proposed intrusion classification scheme may protect smart power grid systems against cyber security attacks.
Hamad Naeem, Farhan Ullah 0001, Gautam Srivastava 0001
Expert Syst. J. Knowl. Eng.2
2025 Efficient malware detection using hybrid approach of transfer learning and generative adversarial examples with image representation
abstract
Abstract Identifying malicious intent within a program, also known as malware, is a critical security task. Many detection systems remain ineffective due to the persistent emergence of zero‐day variants, despite the pervasive use of antivirus tools for malware detection. The application of generative AI in the realm of malware visualization, particularly when binaries are depicted as colour visuals, represents a significant advancement over traditional machine‐learning approaches. Generative AI generates various samples, minimizing the need for specialized knowledge and time‐consuming analysis, hence boosting zero‐day attack detection and mitigation. This paper introduces the Deep Convolutional Generative Adversarial Network for Zero‐Shot Learning (DCGAN‐ZSL), leveraging transfer learning and generative adversarial examples for efficient malware classification. First, a normalization method is proposed, resizing malicious images to 128 × 128 or 300 × 300 for standardized input, enhancing feature transformation for improved malware pattern recognition. Second, greyscale representations are converted into colour images to augment feature extraction, providing a richer input for enhanced model performance in malware classification. Third, a novel DCGAN with progressive training improves model stability, mode collapse, and image quality, thus advancing generative model training. We apply the Attention ResNet‐based transfer learning method to extract texture features from generated samples, which increases security evaluation performance. Finally, the ZSL for zero‐day malware presents a novel method for identifying previously unknown threats, indicating a significant advancement in cybersecurity. The proposed approach is evaluated using two standard datasets, namely dumpware and malimg, achieving malware classification accuracies of 96.21% and 98.91%, respectively.
Yue Zhao 0014, Farhan Ullah 0001, Chien-Ming Chen 0001, Mohammed Amoon, Saru Kumari
Expert Syst. J. Knowl. Eng.2
2025 Measuring student attention based on EEG brain signals using deep reinforcement learning
Asad Ur Rehman, Xiaochuan Shi, Farhan Ullah 0001, Chao Ma 0008
Expert Syst. Appl.3
2025 Elevating e-health excellence with IOTA distributed ledger technology: Sustaining data integrity in next-gen fog-driven systems
Mian Muhammad Waseem Iqbal, Ammar Hassan, Awais Ahmad 0001, Farhan Ullah 0001, Gautam Srivastava 0001
Future Gener. Comput. Syst.5
2025 Homomorphic Encryption Applications for IoT and Light-Weighted Environments: A Review
abstract
Homomorphic encryption (HE) is one of the more sophisticated methods of homomorphic cryptography (HC). HC efficiently contacts the interacting parties in open IoT and light-weighted network environments. This approach is capable of analyzing encrypted data without decryption. The operations use private and public keys. Then, during the assessment or evaluation, users may access the original data. Before conducting tests or evaluations, the customer must first encrypt the data and then decrypt it. Since consumers use several main cycles for the whole operation, which creates noise and computation overheads, the growth rate of computation overheads has increased. The growing ratio of noise to computation rate can interrupt the whole system, resulting in machine instability, protection, and privacy concerns. To resolve the security and privacy issues, the proposed schemes used different hardness assumptions, such as over-integer, learning with error, ideal lattices, bootstrapping, etc. In this article, we presents a comprehensive review of HE and its many varieties. The numerous possible applications of HE are covered at a high level in order to highlight the extent to which HE is used in the IoT and other lighted-weighted intelligent industry environments in a variety of various domains.
Shamsher Ullah, Jianqiang Li 0001, Jie Chen 0027, Ikram Ali, Salabat Khan, Muhammad Tanveer Hussain, Farhan Ullah 0001, Victor C. M. Leung
IEEE Internet Things J.7
2025 EIDS-DTL: Edge-Based Intrusion Detection System for IoUAVs Using Metaheuristic Task Optimization and Deep Transfer Learning
abstract
The integration of Unmanned Aerial Vehicles (UAVs) with the Internet of Things (IoT), also known as IoUAVs, facilitates real-time data transmission and coordinated operations in critical applications such as smart agriculture, disaster response, and infrastructure monitoring. The growing development of IoT has, however, made IoUAVs vulnerable to emerging cyberattacks that could disrupt these essential services. Deep learning can detect hidden attack patterns, but power and processing constraints make it challenging for resource-constrained IoUAVs. Edge computing offloads real-time analysis tasks, but optimizing workloads with unpredictable connectivity and high latency requirements for intrusion detection remains challenging. To address these challenges, this paper proposes a novel Edge-Based Intrusion Detection System (EIDS) that introduces two key innovations. We developed a metaheuristic task optimization technique for the IoUAV edge environment to efficiently manage computational loads and resources. Second, a Deep Transfer Learning (DTL) technique optimized for intrusion detection minimizes training time and computational overhead. Our novel EIDS-DTL technology synergistically incorporates these components for powerful intrusion detection. Our method optimizes feature extraction from IoUAV network traffic by purifying, filtering, and normalizing data. By fine-tuning pre-trained models, the system achieves high accuracy in identifying malicious activity while ensuring optimal performance in resource-constrained environments. Experimental results on two benchmark datasets demonstrate classification accuracies of 98.95% and 99.27%, outperforming existing approaches by up to 5% in accuracy while maintaining high precision, recall, and F1 scores. The proposed method enhances accuracy and efficiency, providing an effective solution for IoUAV security and edge optimization.
Farhan Ullah 0001, Gautam Srivastava 0001, Shamsher Ullah, Leonardo Mostarda, Jawad Ahmad 0001
IEEE Internet Things J.1
2025 TFedSec-HI: Transformer-Driven Federated Security for IoT-Enabled Healthcare Industry 5.0 on Non-IID Data
abstract
The Internet of Things (IoT) enhances the healthcare industry 5.0 by enabling connected devices and data-driven treatments, but it also introduces cyber threats such as data breaches, and unauthorized access. Mobile Edge Computing (MEC) improves security by reducing reliance on cloud transmissions. However, challenges such as Non-Independent and Identically Distributed (Non-IID) data and device intermittency affect security models in healthcare that require real-time analytics and reliable automation. These limitations are crucial in sensitive medical applications that require real-time analytics and reliable automation. This paper proposes TFedSec-HI, a Transformerdriven Federated Learning (TFL) for improving threat detection in the healthcare industry 5.0. Network traffic data is converted to grayscale and multi-channel RGB images using Local Binary Patterns (LBP) and Sobel edge detection. A lightweight mobile Vision Transformer (ViT) is used for effective feature extraction on edge devices, reducing computational load while retaining high performance. The Federated Proximal (FedProx) algorithm is used during the model aggregation phase to address issues with non-IID data distribution, resulting in consistent and effective learning. The global model is then shared with clients for realtime threat classification. The proposed method is evaluated on two real-world datasets, CICIoT2023 and CICIoMT2024, resulting in exceptional classification accuracies of 99.18% and 99.74%, respectively. TFedSec-HI addresses the Non-IID data challenges in Industry 5.0 healthcare by utilizing TFL to enable private, and adaptive threat detection across medical IoT devices.
Yue Zhao 0014, Farhan Ullah 0001, Khalid Mahmood 0002, Jawad Ahmad 0001, Ali Kashif Bashir, Nazik Alturki
IEEE Internet Things J.2
2024 Homomorphic Cryptography Authentication Scheme to Eliminate Machine Tools Gaps in Industry 4.0
Shamsher Ullah, Jianqiang Li 0001, Farhan Ullah 0001, Diletta Cacciagrano, Muhammad Tanveer Hussain, Victor C. M. Leung
AINA (6)3
2024 Collaborative Intrusion Detection System for Intermittent 10 Vs Using Federated Learning and Deep Swarm Particle Optimization
abstract
Intelligent vehicles have significantly influenced the advancement of Intelligent Transportation Systems (ITS). Smart city consumers increasingly depend on vehicular cloud services, highlighting the need for a stronger Internet of Vehicles (IoV s) architecture. Moreover, smart cities deliver high-performance cloud services using multiple technologies, increasing concerns about communication security across entities exchanging indi-vidual requester data. An intelligent privacy-preserving Intrusion Detection System (IDS) is needed to secure IoV data. This work presents a Federated Learning (FL) approach for intermittent IoVs that uses Deep Swarm Particle Optimisation (DSPO) to choose features optimally while protecting user privacy. This approach enables remote IoVs to access shared data securely, ensuring operational confidentiality and privacy. By integrating DPSO with FL, it enhances data analysis and model training for IoV s, optimizing deep learning models for efficient feature selection in secured distributed environments. This cooperative technique not only protects data privacy but also fosters collaboration among IoV devices. We evaluate the proposed method using two standard datasets, namely CICloV2024 and CICEVSE2024. Despite the intermittent nature of IoVs and imbalanced datasets, our approach gives the highest performance.
Farhan Ullah 0001, Gautam Srivastava 0001, Leonardo Mostarda, Diletta Cacciagrano
DSAA1
2024 An expert system for hybrid edge to cloud computational offloading in heterogeneous MEC-MCC environments
Sheharyar Khan, Jiangbin Zheng 0001, Muhammad Irfan 0009, Farhan Ullah 0001, Sohrab Khan
J. Netw. Comput. Appl.4
2024 A Scalable Federated Learning Approach for Collaborative Smart Healthcare Systems With Intermittent Clients Using Medical Imaging
abstract
The healthcare industry is one of the most vulnerable to cybercrime and privacy violations because health data is very sensitive and spread out in many places. Recent confidentiality trends and a rising number of infringements in different sectors make it crucial to implement new methods that protect data privacy while maintaining accuracy and sustainability. Moreover, the intermittent nature of remote clients with imbalanced datasets poses a significant obstacle for decentralized healthcare systems. Federated learning (FL) is a decentralized and privacy-protecting approach to deep learning and machine learning models. In this article, we implement a scalable FL framework for interactive smart healthcare systems with intermittent clients using chest X-ray images. Remote hospitals may have imbalanced datasets with intermittent clients communicating with the FL global server. The data augmentation method is used to balance datasets for local model training. In practice, some clients may leave the training process while others join due to technical or connectivity issues. The proposed method is tested with five to eighteen clients and different testing data sizes to evaluate performance in various situations. The experiments show that the proposed FL approach produces competitive results when dealing with two distinct problems, such as intermittent clients and imbalanced data. These findings would encourage medical institutions to collaborate and use rich private data to quickly develop a powerful patient diagnostic model.
Farhan Ullah 0001, Gautam Srivastava 0001, Shamsher Ullah, Jerry Chun-Wei Lin, Yue Zhao 0014
IEEE J. Biomed. Health Informatics1
2024 NMal-Droid: network-based android malware detection system using transfer learning and CNN-BiGRU ensemble
Farhan Ullah 0001, Shamsher Ullah, Gautam Srivastava 0001, Jerry Chun-Wei Lin, Yue Zhao 0014
Wirel. Networks1
2023 Development of a deep stacked ensemble with process based volatile memory forensics for platform independent malware detection and classification
Hamad Naeem, Olorunjube James Falana, Farhan Ullah 0001
Expert Syst. Appl.4
2023 Android-IoT Malware Classification and Detection Approach Using Deep URL Features Analysis
abstract
Currently, malware attacks pose a high risk to compromise the security of Android-IoT apps. These threats have the potential to steal critical information, causing economic, social, and financial harm. Because of their constant availability on the network, Android apps are easily attacked by URL-based traffic. In this paper, an Android malware classification and detection approach using deep and broad URL feature mining is proposed. This study entails the development of a novel traffic data preprocessing and transformation method that can detect malicious apps using network traffic analysis. The encrypted URL-based traffic is mined to decrypt the transmitted data. To extract the sequenced features, the N-gram analysis method is used, and afterward, the singular value decomposition (SVD) method is utilized to reduce the features while preserving the actual semantics. The latent features are extracted using the latent semantic analysis tool. Finally, CNN-LSTM, a multi-view deep learning approach, is designed for effective malware classification and detection.
Farhan Ullah 0001, Xiaochun Cheng, Leonardo Mostarda, Sohail Jabbar
J. Database Manag.1
2023 Dynamic Resource Allocation Techniques for Wireless Network Data in Elastic Optical Network Applications
Kangcheng Wu, Nasir Jamal, Farhan Ullah 0001
Mob. Networks Appl.4
2023 Video Resources Recommendation for Online Tourism Teaching in Interactive Network
Haitao Shang, Nasir Jamal, Farhan Ullah 0001
Mob. Networks Appl.4
2022 CroLSSim: Cross-language software similarity detector using hybrid approach of LSA-based AST-MDrep features and CNN-LSTM model
abstract
Software similarity in different programming codes is a rapidly evolving field because of its numerous applications in software development, software cloning, software plagiarism, and software forensics. Currently, software researchers and developers search cross-language open-source repositories for similar applications for a variety of reasons, such as reusing programming code, analyzing different implementations, and looking for a better application. However, it is a challenging task because each programming language has a unique syntax and semantic structure. In this paper, a novel tool called Cross-Language Software Similarity (CroLSSim) is designed to detect similar software applications written in different programming codes. First, the Abstract Syntax Tree (AST) features are collected from different programming codes. These are high-quality features that can show the abstract view of each program. Then, Methods Description (MDrep) in combination with AST is used to examine the relationship among different method calls. Second, the Term Frequency Inverse Document Frequency approach is used to retrieve the local and global weights from AST-MDrep features. Third, the Latent Semantic Analysis-based features extraction and selection method is proposed to extract the semantic anchors in reduced dimensional space. Fourth, the Convolution Neural Network (CNN)-based features extraction method is proposed to mine the deep features. Finally, a hybrid deep learning model of CNN-Long-Short-Term Memory is designed to detect semantically similar software applications from these latent variables. The data set contains approximately 9.5K Java, 8.8K C#, and 7.4K C++ software applications obtained from GitHub. The proposed approach outperforms as compared with the state-of-the-art methods.
Farhan Ullah 0001, Muhammad Rashid Naeem, Hamad Naeem, Xiaochun Cheng, Mamoun Alazab
Int. J. Intell. Syst.1
2022 A perspective trend of hyperelliptic curve cryptosystem for lighted weighted environments
Shamsher Ullah, Jiangbin Zheng 0001, Muhammad Tanveer Hussain, Nizamuddin, Farhan Ullah 0001, Muhammad Umar Farooq 0002
J. Inf. Secur. Appl.5
2022 Explainable artificial intelligence approach in combating real-time surveillance of COVID19 pandemic from CT scan and X-ray images using ensemble model
Farhan Ullah 0001, Jihoon Moon, Hamad Naeem, Sohail Jabbar
J. Supercomput.1
2022 IoT-based Cloud Service for Secured Android Markets using PDG-based Deep Learning Classification
abstract
Software piracy is an act of illegal stealing and distributing commercial software either for revenue or identify theft. Pirated applications on Android app stores are harming developers and their users by clone scammers. The scammers usually generate pirated versions of the same applications and publish them in different open-source app stores. There is no centralized system between these app stores to prevent scammers from publishing pirated applications. As most of the app stores are hosted on cloud storage, therefore a cloud-based interaction system can prevent scammers from publishing pirated applications. In this paper, we proposed IoT-based cloud architecture for clone detection using program dependency analysis. First, the newly submitted APK and possible original files are selected from app stores. The APK Extractor and JDEX decompiler extract APK and DEX files for Java source code analysis. The dependency graphs of Java files are generated to extract a set of weighted features. The Stacked-Long Short-Term Memory (S-LSTM) deep learning model is designed to predict possible clones. Experimental results have shown that the proposed approach can achieve an average accuracy of 95.48% among clones from different application stores.
Farhan Ullah 0001, Muhammad Rashid Naeem, Abdullah Bajahzar, Fadi M. Al-Turjman
ACM Trans. Internet Techn.1
2021 Software plagiarism detection in multiprogramming languages using machine learning approach
abstract
Summary The Software plagiarism, which arises the problem of software piracy is a growing major concern nowadays. It is a serious risk to the software industry that gives huge economic damages every year. The customers may develop a modified version of the original software in other types of programming languages. Furthermore, the plagiarism detection in different types of source codes is a challenging task because each source code may have specific syntax rules. In this paper, we proposed a methodology for software plagiarism detection in multiprogramming languages based on machine learning approaches. The Principal Component Analysis (PCA) is applied for features extraction from source codes without losing the actual information. It extracts features by factor analysis and converts the dataset into normalized linear principal components which are further useful for predictions analysis. Then, the multinomial logistic regression model (MLR) is applied to these components to classify the source codes documents based on predictions. It gives the generalization of logistic regression to handle multiclass problems. Further, the predictors' performance in MLR is evaluated by 2 tailed z test. To apply the experiment, the dataset is collected in five different and popular languages, ie, C, C++, Java, C#, and Python. Each programming language taken in two different case studies, ie, binary search and Stack.
Farhan Ullah 0001, Junfeng Wang 0003, Masood Habib, Shehzad Khalid
Concurr. Comput. Pract. Exp.1
2021 An intelligent decision support system for software plagiarism detection in academia
abstract
The act of source code plagiarism is an academic offense that discourages the learning habits of students. Online support is available through which students can hire professional developers to code their regular programming tasks. These facilities make it easier for students to practice plagiarism. First, raw source codes are cleaned from noisy data to extract meaningful codes as the actual logic is more important to the programmers. Second, pre-processing techniques based on tokenization are used to convert filtered codes into meaningful tokens. It breaks the codes into small instances with the number of occurrences known as the frequency. Thirdly, the local and global weighting scheme method is applied to estimate the significance of each feature in an individual or a group of documents. It helps us greatly to zoom in on the importance of each feature of how effective it is for the next phase. Fourth, the single value decomposition method is used to reduce the dimensions of these features by maintaining the actual semantics of the source codes. This technique is used to remove overloaded noise information and collect only those features that are more effective for plagiarism detection. Fifth, the latent semantic analysis (LSA) technique is used to mine the actual semantics of the source codes in the form of latent variables. After that, the LSA features are used as input to cosine similarity to compute the plagiarism among different source codes. To validate the proposed approach, we used the topic modeling approach to group the relevant features into different topics.
Farhan Ullah 0001, Sohail Jabbar, Leonardo Mostarda
Int. J. Intell. Syst.1
2021 A Deep Learning-based Approach for Emotions Classification in Big Corpus of Imbalanced Tweets
abstract
Emotions detection in natural languages is very effective in analyzing the user's mood about a concerned product, news, topic, and so on. However, it is really a challenging task to extract important features from a burst of raw social text, as emotions are subjective with limited fuzzy boundaries. These subjective features can be conveyed in various perceptions and terminologies. In this article, we proposed an IoT-based framework for emotions classification of tweets using a hybrid approach of Term Frequency Inverse Document Frequency (TFIDF) and deep learning model. First, the raw tweets are filtered using the tokenization method for capturing useful features without noisy information. Second, the TFIDF statistical technique is applied to estimate the importance of features locally as well as globally. Third, the Adaptive Synthetic (ADASYN) class balancing technique is applied to solve the imbalance class issue among different classes of emotions. Finally, a deep learning model is designed to predict the emotions with dynamic epoch curves. The proposed methodology is analyzed on two different Twitter emotions datasets. The dynamic epoch curves are shown to show the behavior of test and train data points. It is proved that this methodology outperformed the popular state-of-the-art methods.
Nasir Jamal, Chen Xianqiao, Fadi M. Al-Turjman, Farhan Ullah 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2020 Malware detection in industrial internet of things based on hybrid image visualization and deep learning model
Hamad Naeem, Farhan Ullah 0001, Muhammad Rashid Naeem, Shehzad Khalid, Danish Vasan, Sohail Jabbar, Saqib Saeed
Ad Hoc Networks2
2020 Plagiarism detection in students' programming assignments based on semantics: multimedia e-learning based smart assessment methodology
Farhan Ullah 0001, Junfeng Wang 0003, Sohail Jabbar, Zhiming Wu, Shehzad Khalid
Multim. Tools Appl.1
2017 Semantic Interoperability in Heterogeneous IoT Infrastructure for Healthcare
abstract
Interoperability remains a significant burden to the developers of Internet of Things’ Systems. This is due to the fact that the IoT devices are highly heterogeneous in terms of underlying communication protocols, data formats, and technologies. Secondly due to lack of worldwide acceptable standards, interoperability tools remain limited. In this paper, we proposed an IoT based Semantic Interoperability Model (IoT-SIM) to provide Semantic Interoperability among heterogeneous IoT devices in healthcare domain. Physicians communicate their patients with heterogeneous IoT devices to monitor their current health status. Information between physician and patient is semantically annotated and communicated in a meaningful way. A lightweight model for semantic annotation of data using heterogeneous devices in IoT is proposed to provide annotations for data. Resource Description Framework (RDF) is a semantic web framework that is used to relate things using triples to make it semantically meaningful. RDF annotated patients’ data has made it semantically interoperable. SPARQL query is used to extract records from RDF graph. For simulation of system, we used Tableau, Gruff-6.2.0, and Mysql tools.
Sohail Jabbar, Farhan Ullah 0001, Shehzad Khalid, Murad Khan, Ki Jun Han
Wirel. Commun. Mob. Comput.2