Saad Shafiq

dblp:142/9327 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-5901-1420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2025 Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection Solutions
abstract
Recent research has revealed that the reported results of an emerging body of deep learning-based techniques for detecting software vulnerabilities are not reproducible, either across different datasets or on unseen samples. This paper aims to provide the foundation for properly evaluating the research in this domain. We do so by analyzing prior work and existing vulnerability datasets for the syntactic and semantic features of code that contribute to vulnerability, as well as features that falsely correlate with vulnerability. We provide a novel, uniform representation to capture both sets of features, and use this representation to detect the presence of both vulnerability and spurious features in code. To this end, we design two types of code perturbations: feature preserving perturbations (FPP) ensure that the vulnerability feature remains in a given code sample, while feature eliminating perturbations (FEP) eliminate the feature from the code sample. These perturbations aim to measure the influence of spurious and vulnerability features on the predictions of a given vulnerability detection solution. To evaluate how the two classes of perturbations influence predictions, we conducted a large-scale empirical study on five state-of-the-art DL-based vulnerability detectors. Our study shows that, for vulnerability features, only$\sim 2 \%$of FPPs yield the undesirable effect of a prediction changing among the five detectors on average. However, on average,$\sim 84 \%$of FEPs yield the undesirable effect of retaining the vulnerability predictions. For spurious features, we observed that FPPs yielded a drop in recall up to 29 % for graph-based detectors. We present the reasons underlying these results and suggest strategies for improving DNN-based vulnerability detectors. We provide our perturbation-based evaluation framework as a public resource to enable independent future evaluation of vulnerability detectors.
Satyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad Medvidovic
ICSE3
2024 Toward Improved Deep Learning-based Vulnerability Detection
abstract
Deep learning (DL) has been a common thread across several recent techniques for vulnerability detection. The rise of large, publicly available datasets of vulnerabilities has fueled the learning process underpinning these techniques. While these datasets help the DL-based vulnerability detectors, they also constrain these detectors' predictive abilities. Vulnerabilities in these datasets have to be represented in a certain way, e.g., code lines, functions, or program slices within which the vulnerabilities exist. We refer to this representation as a base unit. The detectors learn how base units can be vulnerable and then predict whether other base units are vulnerable. We have hypothesized that this focus on individual base units harms the ability of the detectors to properly detect those vulnerabilities that span multiple base units (or MBU vulnerabilities). For vulnerabilities such as these, a correct detection occurs when all comprising base units are detected as vulnerable. Verifying how existing techniques perform in detecting all parts of a vulnerability is important to establish their effectiveness for other downstream tasks. To evaluate our hypothesis, we conducted a study focusing on three prominent DL-based detectors: ReVeal, DeepWukong, and LineVul. Our study shows that all three detectors contain MBU vulnerabilities in their respective datasets. Further, we observed significant accuracy drops when detecting these types of vulnerabilities. We present our study and a framework that can be used to help DL-based detectors toward the proper inclusion of MBU vulnerabilities.
Adriana Sejfia, Satyaki Das, Saad Shafiq, Nenad Medvidovic
ICSE3
2024 Exploring Dependencies Among Inconsistencies to Enhance the Consistency Maintenance of Models
abstract
Consistency maintenance is paramount for software engineering, as it improves/guarantees the quality of artifacts (e.g., models) during maintenance and evolution. To perform this maintenance, consistency rules (CR) are commonly defined and applied to evaluate model elements according to desired properties. By empirical studies, it is known that CRs commonly evaluate similar model elements (e.g., multiple CRs checking the consistency of a UML class). Thus, we hypothesize that CRs can be used as a means to identify dependencies among inconsis-tencies and support consistency maintenance tasks. Currently, however, no study investigates to what extent dependencies can be identified and how they can be used to repair inconsistencies. In this paper, we explore dependencies between CRs to identify and group dependent inconsistencies. For that, we define a metamodel that allows dependencies to be expressed. Further-more, we propose a consistency maintenance and dependency analysis mechanism that uses such a metamodel. Additionally, the approach generates repairs for the inconsistencies, considering the groups of dependencies to identify overlapping and conflicting repairs. To evaluate the approach, we conducted an empirical study with 48 UML models and 27 CRs. The results show that our approach identifies dependencies between inconsistencies (46 % of the inconsistencies have dependencies), within a reasonable time, 10ms on average in the worst case. Results also show that dependent inconsistencies can be grouped and used together to identify repairs that are either overlapping (26 % on average) or conflicting (58 % on average).
Luciano Marchezan, Wesley K. G. Assunção, Edvin Herac, Saad Shafiq, Alexander Egyed
SANER4
2024 Balanced knowledge distribution among software development teams - Observations from open- and closed-source software development
abstract
Summary In software development, developer turnover is among the primary reasons for project failures, leading to a great void of knowledge and strain for newcomers. Unfortunately, no established methods exist to measure how the problem domain knowledge is distributed among developers. Awareness of how this knowledge evolves and is owned by key developers in a project helps stakeholders reduce risks caused by turnover. To this end, this paper introduces a novel, realistic representation of problem domain knowledge distribution: the ConceptRealm. To construct the ConceptRealm, we employ a latent Dirichlet allocation model to represent textual features obtained from 300 K issues and 1.3 M comments from 518 open‐source projects. We analyze whether the newly emerged issues and developers share similar concepts or how aligned the individual developers' concepts are with the team over time. We also investigate the impact of leaving developers on the frequency of concepts. Finally, we also evaluate the soundness of our approach on a closed‐source software project, thus allowing the validation of the results from a practical standpoint. We find out that the ConceptRealm can represent the problem domain knowledge within a project and can be utilized to predict the alignment of developers with issues. We also observe that projects exhibit many keepers independent of project maturity and that abruptly leaving keepers correlates with a decline of their core concepts as the remaining developers cannot quickly familiarize themselves with those concepts.
Saad Shafiq, Christoph Mayr-Dorn, Atif Mashkoor, Alexander Egyed
J. Softw. Evol. Process.1
2024 Code smells in pull requests: An exploratory study
abstract
Abstract The quality of a pull request is the primary factor integrators consider for its acceptance or rejection. Code smells indicate sub‐optimal design or implementation choices in the source code that often lead to a fault‐prone outcome, threatening the quality of pull requests. This study explores code smells in 21k pull requests from 25 popular Java projects. We find that both accepted (37%) and rejected (44%) pull requests have code smells, affected mainly by god classes and long methods. Besides, we observe that smelly pull requests are more complex and challenging to understand as they have significantly large sizes, long latency times, more discussion and review comments, and are submitted by contributors with less experience. Our results show that features used in previous studies for pull request acceptance prediction could be potentially employed to predict smell in incoming pull requests. We propose a dynamic approach to predict the presence of such code smells in the newly added pull requests. We evaluate our approach on a dataset of 25 Java projects extracted from GitHub. We further conduct a benchmark study to compare the performance of eight machine learning classifiers. Results of the benchmark study show that XGBoost is the best‐performing classifier for smell prediction.
Muhammad Ilyas Azeem, Saad Shafiq, Atif Mashkoor, Alexander Egyed
Softw. Pract. Exp.2
2021 NLP4IP: Natural Language Processing-based Recommendation Approach for Issues Prioritization
abstract
This paper proposes a recommendation approach for issues (e.g., a story, a bug, or a task) prioritization based on natural language processing, called NLP4IP. The proposed semi-automatic approach takes into account the priority and story points attributes of existing issues defined by the project stakeholders and devises a recommendation model capable of dynamically predicting the rank of newly added or modified issues. NLP4IP was evaluated on 19 projects from 6 repositories employing the JIRA issue tracking software with a total of 29,698 issues. A comprehensive benchmark study was also conducted to compare the performance of various machine learning models. The results of the study showed an average top@3 accuracy of 81% and a mean squared error of 2.2 when evaluated on the validation set. The applicability of the proposed approach is demonstrated in the form of a JIRA plug-in illustrating predictions made by the newly developed machine learning model. The dataset has also been made publicly available in order to support other researchers working in this domain.
Saad Shafiq, Atif Mashkoor, Christoph Mayr-Dorn, Alexander Egyed
SEAA1
2020 Towards Optimal Assembly Line Order Sequencing with Reinforcement Learning: A Case Study
abstract
The new era of Industry 4.0 is leading towards self-learning and adaptable production systems requiring efficient and intelligent decision making. Achieving high production rate in a short span of time, continuous improvement, and better utilization of resources is crucial for such systems. This paper discusses an approach to achieve production optimization by finding optimal sequences of orders, which yield high throughput using reinforcement learning. The feasibility of our approach is evaluated by simulating a plant modelled on a higher level of abstraction taken from a real assembly line. The applicability of the proposed approach is demonstrated in the form of code utilizing the simulation model. The obtained results show promising accuracy of sequences against corresponding throughput during the simulation process.
Saad Shafiq, Christoph Mayr-Dorn, Atif Mashkoor, Alexander Egyed
ETFA1
2019 Communication Patterns of Kanban Teams and Their Impact on Iteration Performance and Quality
abstract
Software development industry is growing rapidly and so are the time and budget constraints getting stringent. After Scrum, the widely adopted agile method, agile practitioners are now shifting towards Kanban due to its effective communication facilitation, transparency and limited work in progress traits. Since, the industry is in transition from scrum to Kanban therefore we don't find many empirical studies yielding results of adopting Kanban. Therefore, in this study we aim to explore more on Kanban teams. Mainly, we aim to find the impact of Kanban team's communication patterns on their iteration performance and quality. The findings revealed that the centralization communication patterns have negative impact on iteration performance and quality of a project. However, small world communication pattern has positive impact on iteration performance and quality of a project.
Saad Shafiq, Irum Inayat, Muhammad Abbas 0002
SEAA1
2013 A framework to rapidly test SDN use-cases and accelerate middlebox applications
abstract
Software-defined networking (SDN) is envisioned to provide a centralized interface to programmatically manage networking elements. However, despite its conceptual simplicity, current switch and SDN architectures have poor performance with little support to innovate and test novel SDN applications. We propose an application extensibility framework that allows researchers to build new SDN applications without requiring modification to the OpenFlow-based plumbing available today. We also employ both hardware and software packet processing capabilities of switching elements by offloading intensive per-packet processing onto the switch processing pipeline using dynamically loadable packet processing modules (PPMs). Our architecture thus allows flexibility in the type of applications alongside high switching performance. We believe that our architecture will unleash the potential of SDN by inspiring the SDN “killer app”. We evaluate our framework using an encryption middlebox application and show a two orders-of-magnitude improvement over an implementation using the current SDN architecture when using hardware offload blocks.
Rajesh Narayanan, Geng Lin, Affan A. Syed, Saad Shafiq, Fahd Gilani
LCN4