Yasir Hussain

dblp:226/9304 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 AI Stock Market Prediction Dashboard
abstract
We present a deployable dashboard for next-day stock prediction that combines a Long Short-Term Memory (LSTM) model with daily news-based sentiment. Using a 20-day sliding window of normalised close and aggregated VADER sentiment, the model forecasts the next-day close for large-cap equities. Across AAPL, MSFT, TSLA, and GOOGL, the system achieved MAE in the$2.5-\unicode{x0024} 5.0$range and directional accuracy up to 58%. In a two-month backtest, a simple rules-based trading bot driven by the model produced a$+2.5 \%$return with a 59% win rate, outperforming naïve and moving-average baselines. Gains were most evident during news-sensitive periods. Limitations include short horizons, lexicon-based sentiment, and simplified execution assumptions. The dashboard exposes forecasts, sentiment summaries, and simulated trades, offering practical utility for exploratory analysis and teaching.
James Roberts, Hoshang Kolivand, Yasir Hussain, Mostafa Tajdini
DeSE3
2025 A Deep Learning Approach to EEG-Based Diagnosis of Cognitive Skills Impairment: Electrode-Level Analysis Insights
abstract
Early identification of cognitive skill impairments is crucial for timely clinical intervention. This study utilizes EEG data to classify cognitive states into three categories: no impairment, mild impairment, and severe impairment. EEG signals from 88 participants were segmented into 10 -second windows and analyzed using a modified EEGNet deep learning architecture, achieving a classification accuracy of 89 %. In addition to classification, statistical analyzes including Welch's t test and Benjamini-Hochberg's FDR correction were used to identify electrodes significantly affected, particularly F3, F4, C3,$\text{C4}, \text{T3}, \text{T4}, \text{P3}, \text{P4}, \text{O1}, \text{O2}, \text{Fz}, \text{Cz}$and Pz, implanted in memory, attention, and executive functions. Topographic brain activation maps highlighted these regional abnormalities, while spectral analysis revealed altered frequency band distributions across impairment levels. Connectivity analysis also showed decreased functional integration between brain regions in mild and severe cases. Combining deep learning, statistical inference, and EEG-based network features presents a robust framework for diagnosing and interpreting cognitive impairments.
Yasir Hussain, Fatima Mudassar, Aamir Saeed Malik
ICTAI1
2025 A context-aware zero trust-based hybrid approach to IoT-based self-driving vehicles security
Izhar Ahmed Khan, Marwa Keshk, Yasir Hussain, Dechang Pi, Bentian Li, Tanzeela Kousar, Bakht Sher Ali
Ad Hoc Networks3
2025 Robust mobile robot path planning via LLM-based dynamic waypoint generation
Muhammad Taha Tariq, Yasir Hussain, Congqing Wang
Expert Syst. Appl.2
2024 An Empirical Study on Python Library Dependency and Conflict Issues
abstract
With the rapid development of open-source communities, code reuse in Python projects is increasingly common. Developers heavily rely on third-party libraries from the Python central repository. They need to write specific configuration scripts with version constraints to ensure the correct version of dependent libraries when building projects. However, existing research focuses on direct dependency libraries and ignores potential dependencies that may exist in other dependency configuration files (e.g., requirements-dev.txt for development dependencies). To fill this gap, we conduct an in-depth comprehensive study to quantify the distribution of direct and potential dependencies, correlation, and classification along with detection tools of dependency conflict issues with 278 top popular Python library projects, which were collected from the prominent dependency tracking system Libraries.io. Specifically, we first investigate the magnitude distribution of dependencies by parsing library source files. Second, we visualize dependencies among Python libraries to determine the correlation. We then classify types of dependency conflicts from the perspective of third-party libraries. Finally, we compare Python dependency conflict detection and resolution tools for researchers and developers. Our findings show that third-party libraries containing dependencies are more common, with 79.1% having at least one dependent library. Moreover, the dependency relationship among Python libraries is intricate, which generates lots of dependency conflict issues such as version conflicts. The main cause of issues is the conflict among third-party libraries, accounting for 60.13%. Our findings can help developers better understand library dependencies and provide them insights on how to better manage them.
Yu Zhou 0010, Yasir Hussain, Wenhua Yang 0001
QRS3
2024 Exploring the Impact of Vocabulary Techniques on Code Completion: A Comparative Approach
abstract
Integrated Development Environments (IDEs) are pivotal in enhancing productivity with features like code completion in modern software development. Recent advancements in Natural Language Processing (NLP) have empowered neural language models for code completion. In this study, we present an extensive investigation of the impact of open and closed vocabulary systems on the task of code completion. Specifically, we compare open and closed vocabulary systems with various vocabulary sizes to observe their impact on code completion performance. We experiment with three different open vocabulary systems: byte pair encoding (BPE), WordPiece and Unigram to compare them with closed-vocabulary systems to analyze their modeling performance. We also conduct experiments with different context sizes to study their impact on code completion performance. We have experimented using various prominent language models, including one from recurrent neural networks and five from transformers. Our results indicate that vocabulary size significantly impacts modeling performance and can artificially boost the accuracy of code completion models, especially in the case of a closed-vocabulary system. Moreover, we find that different vocabulary systems have varying impacts on token coverage, whereas open-vocabulary systems exhibit better token coverage. Our findings offer valuable insights for building effective code completion models, aiding researchers and practitioners in this field.
Yasir Hussain, Yu Zhou 0010, Izhar Ahmed Khan
Int. J. Softw. Eng. Knowl. Eng.1
2024 A Novel Collaborative SRU Network With Dynamic Behaviour Aggregation, Reduced Communication Overhead and Explainable Features
abstract
Leakage and tampering problems in collection and transmission of biomedical data have attracted much attention as these concerns instigates negative impression regarding privacy, security, and reputation of medical networks. This article presents a novel security model that establishes a threat-vector database based on the dynamic behaviours of smart healthcare systems. Then, an improved and privacy-preserved SRU network is designed that aims to alleviate fading gradient issue and enhance the learning process by reducing computational cost. Then, an intelligent federated learning algorithm is deployed to enable multiple healthcare networks to form a collaborative security model in a personalized manner without the loss of privacy. The proposed security method is both parallelizable and computationally effective since the dynamic behaviour aggregation strategy empowers the model to work collaboratively and reduce communication overhead by dynamically adjusting the number of participating clients. Additionally, the visualization of the decision process based on the explainability of features enhances the understanding of security experts by enabling them to comprehend the underlying data evidence and causal reasoning. Compared to existing methods, the proposed security method is capable of thoroughly analyzing and detecting severe security threats with high accuracy, reduce overhead and lower computation cost along with enhanced privacy of biomedical data.
Izhar Ahmed Khan, Muhammad Imran Razzak, Dechang Pi, Umar Zia, Shaharyar Kamal, Yasir Hussain
IEEE J. Biomed. Health Informatics6
2023 Optimized Tokenization Process for Open-Vocabulary Code Completion: An Empirical Study
abstract
Studies have substantiated the efficacy of deep learning-based models in various source code modeling tasks. These models are usually trained on large datasets that are divided into smaller units, known as tokens, utilizing either an open or closed vocabulary system. The selection of a tokenization method can have a profound impact on the number of tokens generated, which in turn can significantly influence the performance of the model. This study investigates the effect of different tokenization methods on source code modeling and proposes an optimized tokenizer to enhance the tokenization performance. The proposed tokenizer employs a hybrid approach that initializes with a global vocabulary based on the most frequent unigrams and incrementally builds an open-vocabulary system. The proposed tokenizer is evaluated against popular tokenization methods such as Closed, Unigram, WordPiece, and BPE tokenizers, as well as tokenizers provided by large pre-trained models such as PolyCoder and CodeGen. The results indicate that the choice of tokenization method can significantly impact the number of sub-tokens generated, which can ultimately influence the modeling performance of a model. Furthermore, our empirical evaluation demonstrates that the proposed tokenizer outperforms other baselines, achieving improved tokenization performance both in terms of a reduced number of sub-tokens and time cost. In conclusion, this study highlights the significance of the choice of tokenization method in source code modeling and the potential for improvement through optimized tokenization techniques.
Yasir Hussain, Yu Zhou 0010, Izhar Ahmed Khan, Nasrullah Khan, Muhammad Zahid Abbas
EASE1
2023 Understanding and Enhancing Issue Prioritization in GitHub
abstract
GitHub has become a prominent platform for open source software development, facilitating collaboration and communication among a diverse group of contributors. Efficient issue tracking is a crucial aspect of managing projects on GitHub, and labels serve as one of the primary mechanisms for issue prioritization, while various other issue features are also utilized by issue handlers for the same purpose. However, in large projects, prioritizing issues remains a challenge, and the efficacy of using labels or other issue features for prioritization is not well understood. To address this knowledge gap, we conduct a comprehensive empirical study that investigates the role of labels in GitHub issue prioritization, examines the influence of various issue features on prioritization, and assesses the performance of different ranking algorithms based on these impactful features. Our study, conducted on a dataset comprising data from over 1.5 million issues across diverse GitHub projects, provides valuable insights for issue handling in open source platforms and offers guidance for future research in this domain. Specifically, the study reveals the limited effectiveness of labels in issue prioritization, highlights the significance of certain issue features in the prioritization process, and compares the performance of various ranking algorithms for issue prioritization to support issue handlers.
Yingying He, Wenhua Yang 0001, Minxue Pan, Yasir Hussain, Yu Zhou 0010
ASE4
2023 Enhancing Code Completion with Implicit Feedback
abstract
Code completion has become an important feature of today’s integrated development environments (IDEs). This task involves predicting the next code token(s) based on its contextual information within the code. However, most existing code completion approaches do not consider users’ feedback during the completion process. In this paper, we propose a framework, EHOPE (Enhance Code Completion with Implicit Feedback)), which exploits LSTM(Long Short-Term Memory) and pre-trained model BERT(Bidirectional Encoder Representation from Transformers) to enhance the performance of token-level code completion. By leveraging users’ feedback information, we train an LSTM model to supplement the recommendation list. In addition, we re-rank the list of recommendations using the pre-trained model BERT, which is fine-tuned with feedback information. Existing token-level code completion tools can be plugged into EHOPE. We choose two representative code completion approaches from different categories: one based on statistical methods and the other based on deep learning. These approaches serve as baselines to showcase the performance improvements of EHOPE, evaluated using Hit@k (Top-k) and MRR(Mean Reciprocal Rank) metrics. Empirical experiments show that the recommendation performance steadily and substantially improves as the feedback data increases compared with the baselines.
Haonan Jin, Yu Zhou 0010, Yasir Hussain
QRS3
2023 Federated-SRUs: A Federated-Simple-Recurrent-Units-Based IDS for Accurate Detection of Cyber Attacks Against IoT-Augmented Industrial Control Systems
abstract
The security of industrial control systems (ICSs) against cyber-attacks is essential in modern era since ICSs are vital constituent of modern societies and smart cities. However, the augmentation of legacy ICS networks with smart computing and networking technologies [such as Internet of Things (IoT)] has intensely enlarged the surface of attacks against these critical infrastructures. This augmentation makes these networks more vulnerable to cyber-attacks and despite the current security solutions, attackers still find ways to proliferate these networks. The intrusion detection system (IDS) is one of the key security aspect to prevent these networks from contemporary cyber-attacks. Therefore, this article proposes a new IDS model named federated-simple recurrent units (SRUs) for the security of IoT-based ICSs. Specifically, the federated-SRUs IDS model uses an improved simple recurrent units architecture to reduce computational cost and alleviate the gradient vanishing issue in recurrent networks. Then, it performs data aggregation through several communication rounds in the federated architecture which allows multiple ICS networks and stakeholders to build a comprehensive IDS model in a privacy-preserving manner. The performance of the federated-SRUs IDS model is validated through experiments using real-world gas pipeline-based ICS network data, which indicates that it is able to accurately detect intrusions in real time without compromising privacy and security. Experiments also verify that the federated-SRUs model outperforms existing state-of-the-art approaches and thus can serve as a viable IDS method in IoT-based ICS networks.
Izhar Ahmed Khan, Dechang Pi, Muhammad Zahid Abbas, Umar Zia, Yasir Hussain, Hatem Soliman
IEEE Internet Things J.5
2023 Boosting source code suggestion with self-supervised Transformer Gated Highway
Yasir Hussain, Yu Zhou 0010, Senzhang Wang
J. Syst. Softw.1
2023 DFF-SC4N: A Deep Federated Defence Framework for Protecting Supply Chain 4.0 Networks
abstract
The management of contemporary communication networks of supply chain (SC) 4.0 is becoming more complex due to the heterogeneity requirements of new devices concerning the integration of the Internet of Things in the legacy industry networks. Hence, it becomes a challenging task to secure networks of SC 4.0 from cyber-attacks and provide a robust and efficient defence framework that can resist sophisticated attacks. Machine learning-based intelligent detection algorithms are often trained at either a centralized or single server, which makes it difficult to train an effective model and also it violates privacy concerns if gathering data from other servers at the edge. Classical machine learning approaches function on the legacy group of data placed on a central or single server, which brands it the least favored choice for supply chain networks, with data privacy issues. To address these problems, this article proposes a federated learning-based efficient detection model named, DFF-SC4N, to proactively identify intrusions from SC 4.0 networks using distributed local data training. DFF-SC4N uses communication rounds in a federated learning manner having gated recurrent units by only sharing the learned parameters and keeps the data intact on local servers. The accuracy of the global model is optimized by an aggregating model, which updates from multiple servers and multiple SC 4.0 networks. Extensive experiments on real industrial network data demonstrate that the DFF-SC4N outperforms both centralized training models and state-of-the-art peer methods in protecting SC 4.0 networks.
Izhar Ahmed Khan, Nour Moustafa, Dechang Pi, Yasir Hussain, Nauman Ali Khan
IEEE Trans. Ind. Informatics4
2022 Enhancing IIoT networks protection: A robust security model for attack detection in Internet Industrial Control Systems
Izhar Ahmed Khan, Marwa Keshk, Dechang Pi, Nasrullah Khan, Yasir Hussain, Hatem Soliman
Ad Hoc Networks5
2022 Learning to transfer knowledge from RDF Graphs with gated recurrent units
abstract
The Internet is a vital part of today’s ecosystem. The speedy evolution of the Internet has brought up practical issues such as the problem of information retrieval. Several methods have been proposed to solve this issue. Such approaches retrieve the information by using SPARQL queries over the Resource Description Framework (RDF) content which requires a precise match concerning the query structure and the RDF content. In this work, we propose a transfer learning-based neural learning method that helps to search RDF graphs to provide probabilistic reasoning between the queries and their results. The problem is formulated as a classification task where RDF graphs are preprocessed to abstract the N-Triples, then encode the abstracted N-triples into a transitional state that is suitable for neural transfer learning. Next, we fine-tune the neural learner to learn the semantic relationships between the N-triples. To validate the proposed approach, we employ ten-fold cross-validation. The results have shown that the anticipated approach is accurate by acquiring the average accuracy, recall, precision, and f-measure. The achieved scores are 97.52%, 96.31%, 98.45%, and 97.37%, respectively, and outperforms the baseline approaches.
Hatem Soliman, Izhar Ahmed Khan, Yasir Hussain
Intell. Data Anal.3
2022 Exploring the Impact of Balanced and Imbalanced Learning in Source Code Suggestion
abstract
Studies have confirmed the robust performance of machine learning classifiers for various source code modeling tasks. In general, machine learning approaches are incapable of handling imbalanced datasets, since they are sensitive to the choice of diverse classes. Therefore, these approaches may lean towards the classes with a large percentage of observations. In this work, we investigate and explore the impact of balanced and imbalanced learning on source code suggestion task otherwise known as code completion, covering a large number of imbalanced classes. We further explore the impact of vocabulary size on modeling performance. First, we provide the essentials to formulate the problem of source code suggestion as a classification task and investigate the level of imbalanced classes. Second, we train the four most adapted neural language models as a baseline to assess the modeling performance. Third, we impose two diverse class balancing techniques, TomekLinks and AllKNN, to balance the datasets and evaluate their impact on the modeling performance. Finally, we trained these models with a weighted imbalanced learning approach and compared the performance with balanced learning approaches. Additionally, we train models by varying the vocabulary size to study their impact. In total, we trained 230 models on 10 real-world software projects and extensively evaluated these models with widely used performance metrics such as Precision, Recall, FScore, mean reciprocal rank (MRR), and Receiver operating characteristics (ROC). Additionally, we employed ANOVA statistical analysis to study the statistical significance and differences between these approaches. This study has demonstrated that the modeling performance decreases during balanced model training, whereas the weighted imbalance training produces comparable results and is more efficient in terms of time cost. Additionally, this study exhibits that a large size of vocabulary does not necessarily improve the modeling performance when out-of-vocabulary predictions are disregarded.
Yasir Hussain, Yu Zhou 0010, Izhar Ahmed Khan
Int. J. Softw. Eng. Knowl. Eng.1
2021 A privacy-conserving framework based intrusion detection method for detecting and recognizing malicious behaviours in cyber-physical power networks
Izhar Ahmed Khan, Dechang Pi, Nasrullah Khan, Zaheer Ullah Khan, Yasir Hussain, Farman Ali 0002
Appl. Intell.5
2021 Improving source code suggestion with code embedding and enhanced convolutional long short-term memory
abstract
Abstract Source code suggestion is the utmost helpful feature in the integrated development environments that helps to quicken software development by suggesting the next possible source code tokens. The source code contains useful semantic information but is ignored or not utilised to its full potential by existing approaches. To improve the performance of source code suggestion, the authors propose a deep semantic net (DeepSN) that makes use of semantic information of the source code. First, DeepSN uses an enhanced hierarchical convolutional neural network combined with code‐embedding to automatically extract the top‐notch features of the source code and to learn useful semantic information. Next, the source code's long and short‐term context dependencies are captured by using long short‐term memory. We extensively evaluated the proposed approach with three baselines on ten real‐world projects and the results are suggesting that the proposed approach surpasses state‐of‐the‐art approaches. On average, DeepSN achieves 7.6% higher accuracy than the best baseline.
Yasir Hussain, Yu Zhou 0010
IET Softw.1
2021 Global Sensitivity Analysis for Fuzzy RDF Data
abstract
The resource description framework (RDF) was adopted by the World Wide Web (W3C) as an essential semantic web standard and the RDF scheme. It accords the hard semantics in the description and wields the crisp metadata. However, it usually produces vague or ambiguous information. Consequently, fuzzy RDF helps deal with such special data by transforming the crisp values into a fuzzy set. A method for analyzing fuzzy RDF data is proposed in this paper. To this end, first, we decompose the RDF into fuzzy RDF variables. Second, we are designing a model for global sensitivity analysis based on the decomposition of fuzzy RDF. It figures out the ambiguities of fuzzy RDF data. The proposed global sensitivity analysis model provides the importance of fuzzy RDF data by considering the response function’s structure and reselects it to a certain degree. A practical tool for sensitivity analysis of fuzzy RDF data has also been implemented based on the proposed model.
Hatem Soliman, Izhar Ahmed Khan, Yasir Hussain
Int. J. Softw. Eng. Knowl. Eng.3
2020 Towards continues code recommendation and implementation system: An Initial Framework
abstract
In the current era, the auto and reliable recommendation system plays a significant role in human life. The code recommender systems are being used in various source code databases to recommend the most suitable source code to the user. While code recommendation, the code analysis concerning 'code quality' and 'code implementation' is important to recommend the most reliable code by considering the objective of the user. The ultimate aim of this research work is to propose a code recommendation and implementation model using the characteristics of DevOps that assist in extracting, analyzing, implementing, and updating the recommender system continuously. The current study presents an initial framework of the proposed code recommender model. The design of the model is based on the data collected through literature review and by conducting an empirical study with experts. We believe that the proposed model will assist the researchers and practitioners to recommend the most secure and suitable source code according to their requirement.
Muhammad Azeem Akbar, Yu Zhou 0010, Yasir Hussain
EASE5
2020 SIOT-RIMM: Towards Secure IOT-Requirement Implementation Maturity Model
abstract
It is very crucial for an organization to encapsulate the requirements in its early stage when they are intending to build a novel system such as the internet of things (IoT), particularly when it comes to capturing privacy and security requirements to gain the public confidence. The proposed research is focused to develop a secure IoT-requirement implementation maturity model (SIOT-RIMM). The proposed model will assist the software development organizations to improve and modify their requirement engineering processes in terms of security and privacy of IoT. The SIOT-RIMM model will be developed based on the existing IoT literature pertaining to security and privacy, industrial empirical study and understanding of the challenges that could negatively influence the implementation of security and privacy in IoT. To develop the maturity levels of SIOT-RIMM, we will consider the concepts of existing maturity models of other software engineering domains. In this preliminary study, 19 challenges were identified using the SLR approach that might have a negative impact on the IoT requirements engineering process. The identified challenges will contribute to the development of SIOT-RIMM maturity levels.
Haibo Hu 0002, Muhammad Azeem Akbar, Yasir Hussain, Ali Mahmoud Baddour
EASE5
2020 Deep Transfer Learning for Source Code Modeling
abstract
In recent years, deep learning models have shown great potential in source code modeling and analysis. Generally, deep learning-based approaches are problem-specific and data-hungry. A challenging issue of these approaches is that they require training from scratch for a different related problem. In this work, we propose a transfer learning-based approach that significantly improves the performance of deep learning-based source code models. In contrast to traditional learning paradigms, transfer learning can transfer the knowledge learned in solving one problem into another related problem. First, we present two recurrent neural network-based models RNN and GRU for the purpose of transfer learning in the domain of source code modeling. Next, via transfer learning, these pre-trained (RNN and GRU) models are used as feature extractors. Then, these extracted features are combined into attention learner for different downstream tasks. The attention learner leverages from the learned knowledge of pre-trained models and fine-tunes them for a specific downstream task. We evaluate the performance of the proposed approach with extensive experiments with the source code suggestion task. The results indicate that the proposed approach outperforms the state-of-the-art models in terms of accuracy, precision, recall and F-measure without training the models from scratch.
Yasir Hussain, Yu Zhou 0010, Senzhang Wang
Int. J. Softw. Eng. Knowl. Eng.1
2020 CodeGRU: Context-aware deep learning with gated recurrent unit for source code modeling
Yasir Hussain, Yu Zhou 0010, Senzhang Wang
Inf. Softw. Technol.1