VLDB 2026 Research / reviewers in the wild / expert
Ashkan Sami
dblp:95/5051
· DBLP profile ↗
31ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-0023-9543ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests
Illia Dovhoshliubnyi, Nima Soroush, Ashkan Sami, Alexander E. I. Brownlee |
SSBSE | 3 |
| 2026 | Secure coding with AI - from detection to repairabstractAbstract While several studies have examined the security of code generated by GPT and other Large Language Models (LLMs), most have relied on controlled experiments rather than real developer interactions. This paper investigates the security of GPT-generated code extracted from the DevGPT dataset and evaluates the ability of current LLMs to detect and repair vulnerabilities in this real-world context. We analysed 2,315 C, C++, and C# code snippets using static scanners combined with manual inspection, identifying 56 vulnerabilities across 48 files. These files were then assessed using GPT-4.1, GPT-5, and Claude Opus 4.1 to determine whether these could identify the security issues and, where applicable, to specify the corresponding Common Weakness Enumeration (CWE) numbers and propose fixes. Manual review and re-scanning of the modified code showed that GPT-4.1, GPT-5, and Claude Opus 4.1 correctly detected 46, 44, and 45 vulnerabilities, and successfully repaired 42, 44, and 43 respectively. A comparison of experiments conducted in October 2024 and September 2025 indicates substantial progress, with overall detection and remediation rates improving from roughly 50% to around 75–80%. We also observe that LLM-generated code is about as likely to contain vulnerabilities as developer-written code, and that LLMs may confidently provide incorrect information, posing risks for less experienced developers. Vladislav Belozerov, Peter J. Barclay, Ashkan Sami |
Empir. Softw. Eng. | 3 |
| 2025 | Exploring the black box: analysing explainable AI challenges and best practices through stack exchange discussionsabstractAbstract Explainable Artificial Intelligence (XAI) is a crucial domain within research and industry, aiming to develop AI models that provide human-understandable explanations for their decisions. While the challenges in AI, deep learning, and big data have been extensively explored, the specific concerns of XAI developers have received limited attention. To address this gap, we analysed discussions on Stack Exchange websites to delve into these issues. Through a combination of automated and Manual analysis, we identified 6 overarching categories, 10 distinct topics, and 40 sub-topics commonly discussed by developers. Our examination revealed a steady rise in discussions on XAI since late 2015, initially focusing on conceptualisation and practical applications, with a notable surge in activity across all topic categories since 2019. Notably, Concepts and Applications, Tools Troubleshooting, and Neural Networks Interpretation emerged as the most popular topics. Troubleshooting challenges were commonly encountered with tools like SHAP, ELI5, and AIF360, while visualisation issues were prevalent with Yellowbrick and SHAP. Furthermore, our analysis suggests that addressing questions related to XAI poses greater difficulty compared to other machine-learning questions. Mohammad Mahdi Sayyadnejad, Ali Asgari, Ashkan Sami, Hooman Tahayori |
Empir. Softw. Eng. | 3 |
| 2025 | Reputation Gaming in Crowd Technical Knowledge SharingabstractStack Overflow incentive system awards users with reputation scores to ensure quality. The decentralized nature of the forum may make the incentive system prone to manipulation. This article offers, for the first time, a comprehensive study of the reported types of reputation manipulation scenarios that might be exercised in Stack Overflow and the prevalence of such reputation gamers by a qualitative study of 1,697 posts from meta Stack Exchange sites. We found four different types of reputation fraud scenarios, such as voting rings where communities form to upvote each other repeatedly on similar posts. We developed algorithms that enable platform managers to automatically identify these suspicious reputation gaming scenarios for review. The first algorithm identifies isolated/semi-isolated communities where probable reputation frauds may occur mostly by collaborating with each other. The second algorithm looks for sudden unusual big jumps in the reputation scores of users. We evaluated the performance of our algorithms by examining the reputation history dashboard of Stack Overflow users from the Stack Overflow Web site. We observed that around 60–80% of users flagged as suspicious by our algorithms experienced reductions in their reputation scores by Stack Overflow. Iren Mazloomzadeh, Gias Uddin 0001, Foutse Khomh, Ashkan Sami |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Investigating Markers and Drivers of Gender Bias in Machine TranslationsabstractImplicit gender bias in Large Language Models (LLMs) is a well-documented problem that needs to be better understood in order to be addressed effectively. Implications of gender introduced into automatic translations can perpetuate real-world biases in Software Engineering and other domains. However, some LLMs use heuristics or post-processing to mask such bias, which makes investigation more difficult. Here, we examine bias in language models via back-translation, using the DeepL online translation service to investigate the bias evinced when repeatedly translating a set of 56 Software Engineering tasks used in a previous study. Each statement starts with ‘she’, and is translated first into a ‘genderless’ intermediate language then back into English; we then examine pronoun-choice in the back-translated texts. We believe this approach provides a useful alternative to large-scale surveys in mapping biases. We expand prior research in the following ways: (1) by comparing results across five intermediate languages, namely Finnish, Indonesian, Estonian, Turkish and Hungarian; (2) by proposing a novel metric for assessing the variation in gender implied in repeated translations of the same phrase, avoiding the over-interpretation of individual pronouns, apparent in earlier work; (3) by investigating sentence features that drive bias; (4) and by comparing results from three time-lapsed datasets to establish the reproducibility of the approach. We found that some languages display similar patterns of pronoun use, falling into three loose groups, but that patterns vary between groups; this underlines the need to work with multiple languages. We also identify the main verb appearing in a sentence as a likely significant driver of implied gender in the translations. Moreover, we see a good level of replicability in the results, and establish that our variation metric proves robust despite an obvious change in the behaviour of the DeepL translation API during the course of the study. These results show that the back-translation method can provide further insights into bias in language models. Peter J. Barclay, Ashkan Sami |
SANER | 2 |
| 2024 | FortisEDoS: A Deep Transfer Learning-Empowered Economical Denial of Sustainability Detection Framework for Cloud-Native Network SlicingabstractNetwork slicing is envisaged as the key to unlocking revenue growth in 5 G and beyond (B5G) networks. However, the dynamic nature of network slicing and the growing sophistication of DDoS attacks rises the menace of reshaping a stealthy DDoS into an Economical Denial of Sustainability (EDoS) attack. EDoS aims at incurring economic damages to service provider due to the increased elastic use of resources. Motivated by the limitations of existing defense solutions, we propose FortisEDoS, a novel framework that aims at enabling elastic B5G services that are impervious to EDoS attacks. FortisEDoS integrates a new deep learning-powered DDoS anomaly detection model, dubbed CG-GRU, that capitalizes on the capabilities of emerging graph and recurrent neural networks in capturing spatio-temporal correlations to accurately discriminate malicious behavior. Furthermore, FortisEDoS leverages transfer learning to effectively defeat EDoS attacks in newly deployed slices by exploiting the knowledge learned in a previously deployed slice. The experimental results demonstrate the superiority of CG-GRU in achieving higher detection performance of more than 92% with lower computation complexity. They show also that transfer learning can yield an attack detection sensitivity of above 91%, while accelerating the training process by at least 61%. Further analysis shows that FortisEDoS exhibits intuitive explainability of its decisions, fostering trust in deep learning-assisted systems. Chafika Benzaid, Tarik Taleb, Ashkan Sami, Othmane Hireche |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | A Deep Transfer Learning-Powered EDoS Detection Mechanism for 5G and Beyond Network SlicingabstractNetwork slicing is recognized as a key enabler for 5G and beyond (B5G) services. However, its dynamic nature and the growing sophistication of DDoS attacks put it at risk of Economical Denial of Sustainability (EDoS) attack, causing economic losses to service provider due to the increased elastic use of resources. Motivated by the limitations of existing solutions, we propose FortisEDoS, a novel framework that aims at enabling EDoS-aware elastic B5G services. FortisEDoS integrates a new deep learning-based DDoS anomaly detection model, called CG-GRU, that leverages the capabilities of emerging graph and recurrent neural networks in capturing spatio-temporal correlations to accurately identify malicious behavior, allowing proactive mitigation of EDoS attacks. Moreover, FortisEDoS uses transfer learning to effectively counteract EDoS attacks in newly deployed slices by leveraging the knowledge acquired in previously deployed slice. The experimental results show the superiority of transfer learning-powered CG-GRU in achieving higher detection performance with lower computation overhead, compared to other baseline methods. Chafika Benzaid, Tarik Taleb, Ashkan Sami, Othmane Hireche |
GLOBECOM | 3 |
| 2023 | A case study of fairness in generated images of Large Language Models for Software Engineering tasksabstractBias in Large Language Models (LLMs) has significant implications. Since they have revolutionized content creation on the web, they can lead to more unfair outcomes, lack of inclusivity, reinforcement of stereotypes and ethical and legal concerns. Notably, OpenAI has recently made claims they have introduced a new technique to ensure that DALL-E-2 generates images of people accurately reflect the diversity of the world’s population. In order to investigate bias within the field of Software Engineering, the study utilized DALL-E-2 image generation to assess 56 tasks related to software engineering. Another objective was to determine the impact of OpenAI’s new measures on the generated images for these specific tasks. Two sets of experiments were conducted. In one set, the tasks were prefixed with the clause "As a Software Engineer," while in the other set, only the tasks themselves were used. The tasks were presented in a gender-neutral manner, and the AI was instructed to generate images for each task 20 times. For a female-dominant task of doing administrative tasks, 40 more images were generated. The study revealed a large gender bias in the 2,280 images generated. For instance, in the subset of experiments with prompts explicitly incorporating the phrase "As a software engineer," only 2% of the generated images portrayed female protagonists. In all the images in this setting, male protagonists were dominant and in 45 tasks 100% of the protagonists were male. Notably, images generated without the prefixed clause only had more female protagonists in ‘provide comments on project milestones’ and ‘provide enhancements’, while other tasks did not exhibit a similar pattern. The findings emphasize unsuitability of implemented guardrails and the importance of further research on LLMs assessments. Further research is needed in LLMs to find out where their guardrails fail so companies can address them properly. Mansour Sami, Ashkan Sami, Peter J. Barclay |
ICSME | 2 |
| 2023 | CoBRA without experts: New paradigm for software development effort estimation using COCOMO metricsabstractAbstract Software development effort estimation (SDEE) is a critical activity in developing software. Accurate effort estimation in the early phases of software design life cycle has important effects on the success of software projects. COCOMO (Constructive Cost Model) is a parametric data‐driven SDEE model whose parameters must be calibrated with an organization's local data for accurate estimation. Such data are scarce for most organizations. On the other hand, CoBRA (Cost estimation, Benchmarking, and Risk Assessment) is one of the powerful hybrid methods that need a small number of local historical data for effort estimation. However, data gathering in CoBRA is time‐consuming and costly. To ease the use of CoBRA, in this paper, we design a methodology that extracts CoBRA‐required data from COCOMO datasets. By the proposed method, data collected for COCOMO would be used in CoBRA. Using CoBRA, a more accurate estimation of the required effort would be achieved with fewer number of historical data than what is required to calibrate the COCOMO model. We apply the proposed method on six well‐known public COCOMO datasets and use them in CoBRA. Obtained results depict an increase in the accuracy of estimations in comparison with other existing methods. Elham Feizpour, Hooman Tahayori, Ashkan Sami |
J. Softw. Evol. Process. | 3 |
| 2022 | Which bugs are missed in code reviews: An empirical study on SmartSHARK datasetabstractIn pull-based development systems, code reviews and pull request comments play important roles in improving code quality. In such systems, reviewers attempt to carefully check a piece of code by different unit tests. Unfortunately, sometimes they miss bugs in their review of pull requests, which lead to quality degradations of the systems. In other words, disastrous consequences occur when bugs are observed after merging the pull requests. The lack of a concrete understanding of these bugs led us to investigate and categorize them. In this research, we try to identify missed bugs in pull requests of SmartSHARK dataset projects. Our contribution is twofold. First, we hypothesized merged pull requests that have code reviews, code review comments,or pull request comments after merging, may have missed bugs after the code review. We considered these merged pull requests as candidate pull requests having missed bugs. Based on our assumption, we obtained 3,261 candidate pull requests from 77 open-source GitHub projects. After two rounds of restrictive manual analysis, we found 187 bugs missed in 173 pull requests. In the first step, we found 224 buggy pull requests containing missed bugs after merging the pull requests. Secondly, we defined and finalized a taxonomy that is appropriate for the bugs that we found and then found the distribution of bug categories after analysing those pull requests all over again. The categories of missed bugs in pull requests and their distributions are: semantic (51.34%), build (15.5%), analysis checks (9.09%), compatibility (7.49%), concurrency (4.28%), configuration (4.28%), GUI (2.14%), API (2.14%), security (2.14%), and memory (1.6%). Fatemeh Khoshnoud, Ali Rezaei Nasab, Zahra Toudeji, Ashkan Sami |
MSR | 4 |
| 2022 | An Empirical Study of C++ Vulnerabilities in Crowd-Sourced Code ExamplesabstractSoftware developers share programming solutions in Q&A sites like Stack Overflow, Stack Exchange, Android forum, and so on. The reuse of crowd-sourced code snippets can facilitate rapid prototyping. However, recent research shows that the shared code snippets may be of low quality and can even contain vulnerabilities. This paper aims to understand the nature and the prevalence of security vulnerabilities in crowd-sourced code examples. To achieve this goal, we investigate security vulnerabilities in the C++ code snippets shared on Stack Overflow over a period of 10 years. In collaborative sessions involving multiple human coders, we manually assessed each code snippet for security vulnerabilities following CWE (Common Weakness Enumeration) guidelines. From the 72,483 reviewed code snippets used in at least one project hosted on GitHub, we found a total of 99 vulnerable code snippets categorized into 31 types. Many of the investigated code snippets are still not corrected on Stack Overflow. The 99 vulnerable code snippets found in Stack Overflow were reused in a total of 2859 GitHub projects. To help improve the quality of code snippets shared on Stack Overflow, we developed a browser extension that allows Stack Overflow users to be notified for vulnerabilities in code snippets when they see them on the platform. Morteza Verdi, Ashkan Sami, Jafar Akhondali, Foutse Khomh, Gias Uddin 0001, Alireza Karami Motlagh |
IEEE Trans. Software Eng. | 2 |
| 2021 | Characterization and Prediction of Questions without Accepted Answers on Stack OverflowabstractA fast and effective approach to obtain information regarding software development problems is to search them to find similar solved problems or post questions on community question answering (CQA) websites. Solving coding problems in a short time is important, so these CQAs have a considerable impact on the software development process. However, if developers do not get their expected answers, the websites will not be useful, and software development time will increase. Stack Overflow is the most popular CQA concerning programming problems. According to its rules, the only sign that shows a question poser has achieved the desired answer is the user's acceptance. In this paper, we investigate unresolved questions, without accepted answers, on Stack Overflow. The number of unresolved questions is increasing. As of August 2019, 47% of Stack Overflow questions were unresolved. In this study, we analyze the effectiveness of various features, including some novel features, to resolve a question. We do not use the features that contain information not present at the time of asking a question, such as answers. To evaluate our features, we deploy several predictive models trained on the features of 18 million questions to predict whether a question will get an accepted answer or not. The results of this study show a significant relationship between our proposed features and getting accepted answers. Finally, we introduce an online tool that predicts whether a question will get an accepted answer or not. Currently, Stack Overflow's users do not receive any feedback on their questions before asking them, so they could carelessly ask unclear, unreadable, or inappropriately tagged questions. By using this tool, they can modify their questions and tags to check the different results of the tool and deliberately improve their questions to get accepted answers. Mohamad Yazdaninia, David Lo 0001, Ashkan Sami |
ICPC | 3 |
| 2021 | How Do Users Answer MATLAB Questions on Q&A Sites? A Case Study on Stack Overflow and MathWorksabstractMATLAB is an engineering programming language with various toolboxes that has a dedicated Question and Answer (Q&A) platform on the MathWorks website, which is similar to Stack Overflow (SO). Moreover, some MATLAB users ask their questions on SO. This paper aims to compare these two Q&A platforms to see what kind of questions are asked and how developers answer these questions in each platform. The result of our analysis on 80,382 MATLAB questions on SO and 266,367 questions on MathWorks show that MATLAB questions on topics ranging from the MATLAB software installation to questions related to programming received high votes and accepted answers on MathWorks. However, the questions about basics of programming such as plots, functions, and variables and questions on converting MATLAB code to other programming languages are very likely to receive answers on SO. Our detailed analysis on SO shows that users answer MATLAB questions with the same rate of the accepted answer as other popular programming languages like Java and Python, but the rate of unanswered questions and questions without an accepted answer for Simulink and the three most popular MATLAB toolboxes -- image processing, signal processing, and computer vision -- are very high. To analyze the evolution of MATLAB questions on SO, we studied 80,382 MATLAB questions using the SOTorrent dataset. The patterns in MATLAB questions' evolution are: 1) Most of the revisions to questions are text-related and not on code snippets. 2) Most of the code-related revisions were performed by the original poster (OP). 3) Non-original posters (Non-OPs) usually revise code snippets' appearance, while OPs usually revise code snippets' content and logic. Mahshid Naghashzadeh, Amir Haghshenas, Ashkan Sami, David Lo 0001 |
SANER | 3 |
| 2021 | Outage Cause Detection in Power Distribution Systems Based on Data MiningabstractRealizing the factors involved in power system outages can be effective in reliability improvement. In this article, we analyze the distribution power network outage data to find dominant factors in occurring vegetation-, animal-, and equipment-related outages. After their integration, real outage, weather, and load as the input data are used to extract associated features. In this article, visualization techniques are initially utilized to show the impact of features on the outage occurrence and then association rule mining is used to find factors correlated with each outage type as well as each other. Association rules are mined using Apriori technique, considering the chi-square and lift index as the measures of interestingness. The outage analyses are also performed for each equipment separately to find the associated rules. The results showing the effectiveness and validity of the proposed method to identify the factors connected with outage occurrences can be used for future planning and the operation schedule of distribution power networks. Mohammad Sadegh Bashkari, Ashkan Sami, Mohammad Rastegar |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | On the use of C# Unsafe Code Context: An Empirical Study of Stack OverflowabstractBackground. C# maintains type safety and security by not allowing direct dangerous pointer arithmetic. To improve performance for special cases, pointer arithmetic is provided via an unsafe context. Programmers can use the C# unsafe keyword to encapsulate a code block, which can use pointer arithmetic. In the Common Language Runtime (CLR), unsafe code is referred to as unverifiable code. It then becomes the responsibility of the programmer to ensure the encapsulated code snippet is not dangerous. Naturally, this raises concern on whether such trust is misused by programmers when they promote the use of C# unsafe context. Aim. We aim to analyze the prevalence and vulnerabilities of share code examples using C# unsafe keyword in Stack Overflow (SO) code sharing platform. Method. By using some regular expressions and manual checks, we extracted C# unsafe code relevant posts from SO and categorized them into some software development scenarios. Results. In the entire SO data dump of September 2018, we find 2,283 C# snippets with the unsafe keyword. Among those posts, 27% of posts are about Image processing, where unsafe codes are mainly used for performance reasons. The second most popular category by 21% of the codes in the posts is used for 'Interoperability' reasons. That is 'unsafe' is used to enable 'Interoperability' between C# managed codes and unmanaged codes. The 'stackalloc' operator is the third category with 9% of unsafe code posts. The stackalloc operator allocates a block of memory on the stack. Since C# 7.2, Microsoft recommends against using 'stackalloc' in unsafe context whenever possible. Manual inspection shows 67 code snippets with dangerous functions that can introduce vulnerability if not used with caution (e.g., buffer overflow). Finally, 35% of 'Interoperability' posts have 'P/Invoke' tag were used outside NativeMethods class, which is in contrast to Microsoft design suggestion. Conclusion. Our study leads to 7 main findings, and these findings show the importance of cautiously using this feature. Ehsan Firouzi, Ashkan Sami, Foutse Khomh, Gias Uddin 0001 |
ESEM | 2 |
| 2018 | Mining and extraction of personal software process measures through IDE interaction logsabstractThe Personal Software Process (PSP) is an effective software process improvement method that heavily relies on manual collection of software development data. This paper describes a semi-automated method that reduces the burden of PSP data collection by extracting the required time and size of PSP measurements from IDE interaction logs. The tool mines enriched event data streams so can be easily generalized to other developing environment also. In addition, the proposed method is adaptable to phase definition changes and creates activity visualizations and summarizations that are helpful for software project management. Tools and processed data used for this paper are available on GitHub at: https://github.com/unknowngithubuser1/data. Alireza Joonbakhsh, Ashkan Sami |
MSR | 2 |
| 2017 | MAAR: Robust features to detect malicious activity based on API calls, their arguments and return values
Zahra Salehi, Ashkan Sami, Mahboobeh Ghiasi |
Eng. Appl. Artif. Intell. | 2 |
| 2017 | A statistical unsupervised method against false data injection attacks: A visualization-based approach
Mostafa Mohammadpourfard, Ashkan Sami, Ali Reza Seifi |
Expert Syst. Appl. | 2 |
| 2016 | Employing secure coding practices into industrial applications: a case study
Abdullah Khalili, Ashkan Sami, Mahdi Azimi, Sara Moshtari, Zahra Salehi, Mahboobeh Ghiasi, Ali Akbar Safavi |
Empir. Softw. Eng. | 2 |
| 2016 | CIP-UQIM: A unified model for quality improvement in software SME's based on CMMI level 2 and 3
Hosein Rahmani, Ashkan Sami, Abdullah Khalili |
Inf. Softw. Technol. | 2 |
| 2015 | Dynamic VSA: a framework for malware detection based on register contents
Mahboobeh Ghiasi, Ashkan Sami, Zahra Salehi |
Eng. Appl. Artif. Intell. | 2 |
| 2015 | Scaling up the hybrid Particle Swarm Optimization algorithm for nominal data-setsabstractRecently, the hybrid Particle Swarm Optimisation/Ant Colony Optimisation (PSO/ACO) has been proposed for discovery of classification rules. An improved version of this hybrid scheme, PSO/ACO2 algorithm, can directly cope with nominal attributes without converting them into numerical ones. Although PSO/ACO2 can handle nominal values, it suffers from high computational complexity for large datasets. Beside variety of classification methods which exist to provide more compact set of rules, this study propose an approach which reduces the computational complexity of PSO/ACO2 in order to make it suitable for classification of large datasets. This work is developed the K-mode as a method of sampling from the datasets. In this regard a modification is employed to this algorithm in order to decrease the computational time as well as maintaining the accuracy of the algorithm. Further contribution of this paper is utilizing a new fitness measure for the algorithm. This measure has a robust theoretical background. The combination of the proposed modified K-mode method and the introduced fitness measure led to speed the obtained results up. The experimental result shows the efficiency of the proposed algorithm in comparison with its competitors. Hedayatollah Dallaki, Kimia Bazargan Lari, Ali Hamzeh, Sattar Hashemi, Ashkan Sami |
Intell. Data Anal. | 5 |
| 2015 | Knowledge discovery and sequence-based prediction of pandemic influenza using an integrated classification and association rule mining (CBA) algorithmabstractPandemic influenza is a major concern worldwide. Availability of advanced technologies and the nucleotide sequences of a large number of pandemic and non-pandemic influenza viruses in 2009 provide a great opportunity to investigate the underlying rules of pandemic induction through data mining tools. Here, for the first time, an integrated classification and association rule mining algorithm (CBA) was used to discover the rules underpinning alteration of non-pandemic sequences to pandemic ones. We hypothesized that the extracted rules can lead to the development of an efficient expert system for prediction of influenza pandemics. To this end, we used a large dataset containing 5373 HA (hemagglutinin) segments of the 2009 H1N1 pandemic and non-pandemic influenza sequences. The analysis was carried out for both nucleotide and protein sequences. We found a number of new rules which potentially present the undiscovered antigenic sites at influenza structure. At the nucleotide level, alteration of thymine (T) at position 260 was the key discriminating feature in distinguishing non-pandemic from pandemic sequences. At the protein level, rules including I233K, M334L were the differentiating features. CBA efficiently classifies pandemic and non-pandemic sequences with high accuracy at both the nucleotide and protein level. Finding hotspots in influenza sequences is a significant finding as they represent the regions with low antibody reactivity. We argue that the virus breaks host immunity response by mutation at these spots. Based on the discovered rules, we developed the software, "Prediction of Pandemic Influenza" for discrimination of pandemic from non-pandemic sequences. This study opens a new vista in discovery of association rules between mutation points during evolution of pandemic influenza. Fatemeh Kargarfard, Ashkan Sami, Esmaeil Ebrahimie |
J. Biomed. Informatics | 2 |
| 2015 | SePaS: Word sense disambiguation by sequential patterns in sentencesabstractAbstract An open problem in natural language processing is word sense disambiguation (WSD). A word may have several meanings, but WSD is the task of selecting the correct sense of a polysemous word based on its context. Proposed solutions are based on supervised and unsupervised learning methods. The majority of researchers in the area focused on choosing proper size of ‘n’ in n-gram that is used for WSD problem. In this research, the concept has been taken to a new level by using variable ‘n’ and variable size window. The concept is based on the iterative patterns extracted from the text. We show that this type of sequential pattern is more effective than many other solutions for WSD. Using regular data mining algorithms on the extracted features, we significantly outperformed most monolingual WSD solutions. The state-of-the-art results were obtained using external knowledge like various translations of the same sentence. Our method improved the accuracy of the multilingual system more than 4 percent, although we were using monolingual features. Masoud Narouei, Mansour Ahmadi, Ashkan Sami |
Nat. Lang. Eng. | 3 |
| 2015 | DLLMiner: structural mining for malware detectionabstractAbstract Existing anti‐malware products usually use signature‐based techniques as their main detection engine. Although these methods are very fast, they are unable to provide effective protection against newly discovered malware or mutated variant of old malware. Heuristic approaches are the next generation of detection techniques to mitigate the problem. These approaches aim to improve the detection rate by extracting more behavioral characteristics of malware. Although these approaches cover the disadvantages of signature‐based techniques, they usually have a high false positive, and evasion is still possible from these approaches. In this paper, we propose an effective and efficient heuristic technique based on static analysis that not only detect malware with a very high accuracy, but also is robust against common evasion techniques such as junk injection and packing. Our proposed system is able to extract behavioral features from a unique structure in portable executable, which is called dynamic‐link library dependency tree, without actually executing the application. Copyright © 2015 John Wiley & Sons, Ltd. Masoud Narouei, Mansour Ahmadi, Giorgio Giacinto, Hassan Takabi, Ashkan Sami |
Secur. Commun. Networks | 5 |
| 2014 | CBR Clone Based Software Flaw Detection IssuesabstractThe biggest problem in computer security is that most systems aren't constructed with security in mind. Being aware of common security weaknesses in programming might sound like a good way to avoid them, but awareness by itself often proves to be insufficient. Understanding security is one thing and applying your understanding in a complete and consistent fashion to meet your security goals is quite another. For this reason, static analysis is advocated as a technique for finding common security errors in source code. Manual security static analysis is a tedious work, so automatic tools which can guide programmers to detect security concerns is suggested. Good static analysis tools provide a fast way to get a detailed security related evaluation of program code. In this paper a new architecture (CBRFD) for software flaw detector, based on the concept of clone detection and case base reasoning, is proposed and various issues which concern detection of security weakness of codes through code clone detector is investigated. Ali Reza Honarvar, Ashkan Sami |
SIN | 2 |
| 2014 | Entropy-based outlier detection using semi-supervised approach with few positive examples
Armin Daneshpazhouh, Ashkan Sami |
Pattern Recognit. Lett. | 2 |
| 2009 | A Recursive Classifier System for Partially Observable EnvironmentsabstractPreviously we introduced Parallel Specialized XCS (PSXCS), a distributed-architecture classifier system that detects aliased environmental states and assigns their handling to created subordinate XCS classifier systems. PSXCS uses a history-window approach, but with novel efficiency since the subordinateXCSs, which employ the windows, are only spawned for parts of the state space that are actually aliased. However, because the window lengths are finite and set manually, PSXCS may fail to be optimal in difficult test mazes. This paper introduces Recursive PSXCS (RPSXCS) that automatically spawns windows wherever more history is required. Experimental results show that RPSXCS is both more powerful and learns faster than PSXCS. The present research suggests new potential for history approaches to partially observable environments. Ali Hamzeh, Sattar Hashemi, Ashkan Sami, Adel Torkaman Rahmani |
Fundam. Informaticae | 3 |
| 2009 | Agent Based Decision Tree Learning: a Novel ApproachabstractDecision trees are one of the most effective and widely used induction methods that have received a great deal of attention over the past twenty years. When decision tree induction algorithms were used with uncertain rather than deterministic data, the result is a complete tree, which can classify most of the unseen samples correctly. This tree would be pruned in order to reduce its classification error and over-fitting. Recently, multi agent researchers concentrated on learning from large databases. In this paper we present a novel multi agent learning method that is able to induce a decision tree from distributed training sets. Our method is based on combination of separate decision trees each provided by one agent. Hence an agent is provided to aggregate results of the other agents and induces the final tree. Our empirical results suggest that the proposed method can provide significant benefits to distributed data classification. Mohsen Rahmani, Sattar Hashemi, Ali Hamzeh, Ashkan Sami |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2006 | Obstacles and Misunderstandings Facing Medical Data Mining
Ashkan Sami |
ADMA | 1 |
| 2006 | OSDM: Optimized Shape Distribution Method
Ashkan Sami, Ryoichi Nagatomi, Makoto Takahashi, Takeshi Tokuyama |
ADMA | 1 |