Cagatay Catal

dblp:88/2962 · also Çagatay Çatal · DBLP profile ↗
← Back
41ranked-venue papers
14as first author
25since 2021 · last 2025
0000-0003-0959-2930ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 14 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Food fraud detection using explainable artificial intelligence
abstract
Abstract Recently, the global food supply chain has become increasingly complex, and its scalability has grown. From farm to fork, the performance of food‐producing systems is influenced by significant changes in the environment, population and economy. These changes may cause an increase in food fraud and safety hazards and hence, harm human health. Adopting artificial intelligence (AI) technology in the food supply chain is one strategy to reduce these hazards. Although the use of AI has been rising in numerous industries, such as precision nutrition, self‐driving cars, precision agriculture, precision medicine and food safety, much of what AI systems do is a black box due to its poor explainability. This study covers numerous use cases of food fraud risk prediction using explainable artificial intelligence (XAI) techniques, such as LIME, SHAP and WIT. We aimed to interpret the predictions of a machine learning model with the aid of these technologies. The case study was performed on a food fraud dataset using adulteration/fraud notifications retrieved from the Rapid Alert System for Food and Feed system and economically motivated adulteration database. A deep learning model was built based on this dataset and XAI tools have been investigated on the proposed deep learning model. Both features and shortcomings of the current XAI tools in the food fraud area have been presented.
Okan Buyuktepe, Cagatay Catal, Gorkem Kar, Yamine Bouzembrak, Hans J. P. Marvin, Anand K. Gavai
Expert Syst. J. Knowl. Eng.2
2025 Dynamic warping as a sensor reconstruction method for remaining useful life estimation
abstract
This study proposes Dynamic Warping (DW) as a sensor reconstruction method for Remaining Useful Life (RUL) estimation. The method utilizes the DW model for sensor reconstruction, where Median Absolute Deviation measures the reconstruction error, which is expected to increase when abnormal system behavior is measured. We apply an exponential model to the reconstruction error to estimate a system’s RUL. The DW model is based on the Dynamic Time Warping algorithm applied to a non-temporal context. The concept is to preprocess sensor data into a non-temporal motion profile representing a cycle. We validate our proposed DW model with two baseline models: Singular Value Decomposition (SVD) and LSTM Autoencoder (LSTM-AE). The SVD model is applied to the non-temporal motion profile, while the LSTM-AE model is applied to the original sensor data. A case study was conducted at a semiconductor Original Equipment Manufacturer, whose dataset contained information on a bearing failure in a water-cooled direct drive rotary motor. The failure occurred due to increased friction caused by bearing wear. The valuable motion control signal found was torque applied to the shaft for the R and S phases. It was demonstrated that the proposed method is most efficient, and an alarm can be raised 11 hours before failure, after which the RUL can be estimated, which is promising for warning service engineers for this industrial application. This research shows that the DW model could predict maintenance furthest in advance while only needing a fraction of the training data.
Raymon van Dinter, Philippe Leduc, Bedir Tekinerdogan, Cagatay Catal, Yiping Sun
Knowl. Based Syst.5
2024 Design and implementation of a deep learning-empowered m-Health application
abstract
Abstract Many people are unaware of the severity of melanoma disease even though such a disease can be fatal if not treated early. This research aims to facilitate the diagnosis of melanoma disease in people using a mobile health application because some people do not prefer to visit a dermatologist due to several concerns such as feeling uncomfortable by exposing their bodies. As such, a skincare application was developed so that a user can easily analyze a mole at any part of the body and get the diagnosis results quickly. In the first phase, the corresponding image is extracted and sent to a web service. Later, the web service classifies using the pre-trained model built based on a deep learning algorithm. The final phase displays the confidence rates on the mobile application. The proposed model utilizes the Convolutional Neural Network and provides 84% accuracy and 72% precision. The results demonstrate that the proposed model and the corresponding mobile application provide remarkable results for addressing the specified health problem.
Akhan Akbulut, Sara Desouki, Sara AbdelKhaliq, Layal Khantomani, Cagatay Catal
Multim. Tools Appl.5
2024 Hierarchical multi-head attention LSTM for polyphonic symbolic melody generation
abstract
Abstract Creating symbolic melodies with machine learning is challenging because it requires an understanding of musical structure and the handling of inter-dependencies and long-term dependencies. Learning the relationship between events that occur far apart in time in music poses a considerable challenge for machine learning models. Another notable feature of music is that notes must account for several inter-dependencies, including melodic, harmonic, and rhythmic aspects. Baseline methods, such as RNNs, LSTMs, and GRUs, often struggle to capture these dependencies, resulting in the generation of musically incoherent or repetitive melodies. As such, in this study, a hierarchical multi-head attention LSTM model is proposed for creating polyphonic symbolic melodies. This enables our model to generate more complex and expressive melodies than previous methods, while still being musically coherent. The model allows learning of long-term dependencies at different levels of abstraction, while retaining the ability to form inter-dependencies. The study has been conducted on two major symbolic music datasets, MAESTRO and Classical-Music MIDI, which feature musical content encoded on MIDI. The artistic nature of music poses a challenge to evaluating the generated content and qualitative analysis are often not enough. Thus, human listening tests are conducted to strengthen the evaluation. Qualitative analysis conducted on the generated melodies shows significantly improved loss scores on MSE over baseline methods, and is able to generate melodies that were both musically coherent and expressive. The listening tests conducted using Likert-scale support the qualitative results and provide better statistical scores over baseline methods.
Ahmet Kasif, Selcuk Sevgen, Alper Ozcan, Cagatay Catal
Multim. Tools Appl.4
2024 The use of multi-task learning in cybersecurity applications: a systematic literature review
abstract
Abstract Cybersecurity is crucial in today’s interconnected world, as digital technologies are increasingly used in various sectors. The risk of cyberattacks targeting financial, military, and political systems has increased due to the wide use of technology. Cybersecurity has become vital in information technology, with data protection being a major priority. Despite government and corporate efforts, cybersecurity remains a significant concern. The application of multi-task learning (MTL) in cybersecurity is a promising solution, allowing security systems to simultaneously address various tasks and adapt in real-time to emerging threats. While researchers have applied MTL techniques for different purposes, a systematic overview of the state-of-the-art on the role of MTL in cybersecurity is lacking. Therefore, we carried out a systematic literature review (SLR) on the use of MTL in cybersecurity applications and explored its potential applications and effectiveness in developing security measures. Five critical applications, such as network intrusion detection and malware detection, were identified, and several tasks used in these applications were observed. Most of the studies used supervised learning algorithms, and there were very limited studies that focused on other types of machine learning. This paper outlines various models utilized in the context of multi-task learning within cybersecurity and presents several challenges in this field.
Shimaa Ibrahim, Cagatay Catal, Thabet Kacem
Neural Comput. Appl.2
2024 Stacking-based ensemble learning for remaining useful life estimation
abstract
Abstract Excessive and untimely maintenance prompts economic losses and unnecessary workload. Therefore, predictive maintenance models are developed to estimate the right time for maintenance. In this study, predictive models that estimate the remaining useful life of turbofan engines have been developed using deep learning algorithms on NASA’s turbofan engine degradation simulation dataset. Before equipment failure, the proposed model presents an estimated timeline for maintenance. The experimental studies demonstrated that the stacking ensemble learning and the convolutional neural network (CNN) methods are superior to the other investigated methods. While the convolution neural network (CNN) method was superior to the other investigated methods with an accuracy of 93.93%, the stacking ensemble learning method provided the best result with an accuracy of 95.72%.
Begum Ay Ture, Akhan Akbulut, Cagatay Catal
Soft Comput.4
2024 Privacy and Integrity Protection for IoT Multimodal Data Using Machine Learning and Blockchain
abstract
With the wide application of Internet of Things (IoT) technology, large volumes of multimodal data are collected and analyzed for various diagnoses, analyses, and predictions to help in decision-making and management. However, the research on protecting data integrity and privacy is quite limited, while the lack of proper protection for sensitive data may have significant impacts on the benefits and gains of data owners. In this research, we propose a protection solution for data integrity and privacy. Specifically, our system protects data integrity through distributed systems and blockchain technology. Meanwhile, our system guarantees data privacy using differential privacy and Machine Learning (ML) techniques. Our system aims to maintain the usability of the data for further data analytical tasks of data users, while encrypting the data according to the requirements of data owners. We implement our solution with smart contracts, distributed file systems, and ML models. The experimental results show that our proposed solution can effectively encrypt source IoT data according to the requirements of data users while data integrity can be protected under the blockchain.
Qingzhi Liu, Chenglu Jin, Xiaohan Zhou, Ying Mao 0001, Cagatay Catal, Long Cheng 0003
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Analyzing the performance of long short-term memory architectures for malware detection models
abstract
Summary Malicious software forms a threat to many software‐intensive systems and as such several malware detection approaches have been introduced, often based on sequential data analysis. Long short‐term memory (LSTM) is an artificial recurrent neural network (RNN) architecture that is effective for sequential data analysis, however, no study has yet analyzed the performance of different LSTM architectures for the application of malware detection. In this article, we aim to evaluate and benchmark the performance of LSTM‐based malware detection approaches on specific LSTM architectures to provide insight into malware detection. Our method builds LSTM‐based malware prediction models and performs experiments using different LSTM architectures including Vanilla LSTM, stacked LSTM, bi‐directional LSTM, and CNN‐LSTM. We evaluated the performance of each of these architectures and different configurations. Our study, as a contribution, shows that Bidirectional LSTM with hyperparameter optimization is found to be overperforming other selected LSTM architectures. This study shows that different LSTM approaches and architectures are applicable to the malware detection problem. Quality attributes such as efficiency and accuracy, and the software system architecture adopted for the implementation impact the selection of the LSTM approach.
Çigdem Avci Salma, Bedir Tekinerdogan, Cagatay Catal
Concurr. Comput. Pract. Exp.3
2023 The automation of the development of classification models and improvement of model quality using feature engineering techniques
abstract
Recently pipelines of machine learning-based classification models have become important to codify, orchestrate, and automate the workflow to produce an effective machine learning model. In this article, we propose a framework that combines feature engineering techniques such as data imputation, transformation, and class balancing to compare the performance of different prediction models and select the best final model based on predefined parameters. The proposed framework is extendable and configurable by adding algorithms supported by the CARET package implemented in the R programming language. This framework can generate different machine learning models, which provide comparable results compared to other studies. The framework allows practitioners and researchers to automatically generate different classification models. This research used High-Resolution Orbitrap-based Mass Spectrometers (HRMS) data to create automated prediction models for the first time in literature. We demonstrated the applicability of feature engineering techniques such as data imputation, transformation (e.g., scaling, centering, etc.), and data balancing using several case studies and the proposed semi-automated framework. We showed how the initial prediction models can be improved using the proposed framework.
Sjoerd Boeschoten, Cagatay Catal, Bedir Tekinerdogan, Arjen Lommen, Marco Blokland
Expert Syst. Appl.2
2023 The role of Reinforcement Learning in software testing
abstract
Software testing is applied to validate the behavior of the software system and identify flaws and bugs. Different machine learning technique types such as supervised and unsupervised learning were utilized in software testing. However, for some complex software testing scenarios, neither supervised nor unsupervised machine learning techniques were adequate. As such, researchers applied Reinforcement Learning (RL) techniques in some cases. However, a systematic overview of the state-of-the-art on the role of reinforcement learning in software testing is lacking. The objective of this study is to determine how and to what extent RL was used in software testing. In this study, a Systematic Literature Review (SLR) was conducted on the use of RL in software testing, and 40 primary studies were investigated. This study highlights different software testing types to which RL has been applied, commonly used RL algorithms and architecture for learning, challenges faced, advantages and disadvantages of using RL, and the performance comparison of RL-based models against other techniques. RL has been widely used in software testing but has almost narrowed to two applications. There is a shortage of papers using advanced RL techniques in addition to multi-agent RL. Several challenges were presented in this study.
Amr Abo-eleneen, Ahammed Palliyali, Cagatay Catal
Inf. Softw. Technol.3
2023 Analysis of facial emotion expression in eating occasions using deep learning
abstract
Eating is experienced as an emotional social activity in any culture. There are factors that influence the emotions felt during food consumption. The emotion felt while eating has a significant impact on our lives and affects different health conditions such as obesity. In addition, investigating the emotion during food consumption is considered a multidisciplinary problem ranging from neuroscience to anatomy. In this study, we focus on evaluating the emotional experience of different participants during eating activities and aim to analyze them automatically using deep learning models. We propose a facial expression-based prediction model to eliminate user bias in questionnaire-based assessment systems and to minimize false entries to the system. We measured the neural, behavioral, and physical manifestations of emotions with a mobile app and recognize emotional experiences from facial expressions. In this research, we used three different situations to test whether there could be any factor other than the food that could affect a person's mood. We asked users to watch videos, listen to music or do nothing while eating. This way we found out that not only food but also external factors play a role in emotional change. We employed three Convolutional Neural Network (CNN) architectures, fine-tuned VGG16, and Deepface to recognize emotional responses during eating. The experimental results demonstrated that the fine-tuned VGG16 provides remarkable results with an overall accuracy of 77.68% for recognizing the four emotions. This system is an alternative to today's survey-based restaurant and food evaluation systems.
Elif Yildirim, Fatma Patlar Akbulut, Cagatay Catal
Multim. Tools Appl.3
2023 The use of generative adversarial networks in medical image augmentation
abstract
Abstract Generative Adversarial Networks (GANs) have been widely applied in various domains, including medical image analysis. GANs have been utilized in classification and segmentation tasks, aiding in the detection and diagnosis of diseases and disorders. However, medical image datasets often suffer from insufficiency and imbalanced class distributions. To overcome these limitations, researchers have employed GANs to generate augmented medical images, effectively expanding datasets and balancing class distributions. This review follows the PRISMA guidelines and systematically collects peer-reviewed articles on the development of GAN-based augmentation models. Automated searches were conducted on electronic databases such as IEEE, Scopus, Science Direct, and PubMed, along with forward and backward snowballing. Out of numerous articles, 52 relevant ones published between 2018 and February 2022 were identified. The gathered information was synthesized to determine common GAN architectures, medical image modalities, body organs of interest, augmentation tasks, and evaluation metrics employed to assess model performance. Results indicated that cGAN and DCGAN were the most popular GAN architectures in the reviewed studies. Medical image modalities such as MRI, CT, X-ray, and ultrasound, along with body organs like the brain, chest, breast, and lung, were frequently used. Furthermore, the developed models were evaluated, and potential challenges and future directions for GAN-based medical image augmentation were discussed. This review presents a comprehensive overview of the current state-of-the-art in GAN-based medical image augmentation and emphasizes the potential advantages and challenges associated with GAN utilization in this domain.
Ahmed Makhlouf, Marina Maayah, Nada Abughanam, Cagatay Catal
Neural Comput. Appl.4
2023 A hybrid DNN-LSTM model for detecting phishing URLs
Alper Ozcan, Cagatay Catal, Emrah Donmez, Behcet Senturk
Neural Comput. Appl.2
2023 Metamorphic relation automation: Rationale, challenges, and solution directions
abstract
Abstract Metamorphic testing addresses the issue of the oracle problem by comparing results transformation from multiple test executions. The relationship that governs the output transformation is called metamorphic relation. Metamorphic relations require expert knowledge and the generation of them is considered a time‐consuming task. Researchers have proposed various techniques to automate metamorphic testing, generation, and selection. Although there are several research articles on this issue, there is a lack of overview of the state‐of‐the‐art of metamorphic relation automation. As such, we performed a systematic literature review study to collect, extract, and synthesize the required data. Based on our research questions, the literature was categorized and summarized into different categories. We found that the automation of metamorphic relation is most effective in mathematical and scientific applications. We concluded that some approaches involve analysis of different forms of software‐related information such as control flow graph and program dependence graph as well as an initial set of metamorphic relations. On the other hand, other methods involve analysis of executions of the software functions with random and specific inputs. The results show that this field is still in its infancy with opportunities for novel work, especially in methods utilizing machine learning.
Emran Altamimi, Abdullah Elkawakjy, Cagatay Catal
J. Softw. Evol. Process.3
2023 Just-in-time defect prediction for mobile applications: using shallow or deep learning?
abstract
Abstract Just-in-time defect prediction (JITDP) research is increasingly focused on program changes instead of complete program modules within the context of continuous integration and continuous testing paradigm. Traditional machine learning-based defect prediction models have been built since the early 2000s, and recently, deep learning-based models have been designed and implemented. While deep learning (DL) algorithms can provide state-of-the-art performance in many application domains, they should be carefully selected and designed for a software engineering problem. In this research, we evaluate the performance of traditional machine learning algorithms and data sampling techniques for JITDP problems and compare the model performance with the performance of a DL-based prediction model. Experimental results demonstrated that DL algorithms leveraging sampling methods perform significantly worse than the decision tree-based ensemble method. The XGBoost-based model appears to be 116 times faster than the multilayer perceptron-based (MLP) prediction model. This study indicates that DL-based models are not always the optimal solution for software defect prediction, and thus, shallow, traditional machine learning can be preferred because of better performance in terms of accuracy and time parameters.
Raymon van Dinter, Cagatay Catal, Görkem Giray, Bedir Tekinerdogan
Softw. Qual. J.2
2023 Temporal fusion transformer-based prediction in aquaponics
abstract
Abstract Aquaponics offers a soilless farming ecosystem by merging modern hydroponics with aquaculture. The fish food is provided to the aquaculture, and the ammonia generated by the fish is converted to nitrate using specialized bacteria, which is an essential resource for vegetation. Fluctuations in the ammonia levels affect the generated nitrate levels and influence farm yields. The sensor-based autonomous control of aquaponics can offer a highly rewarding solution, which can enable much more efficient ecosystems. Also, manual control of the whole aquaponics operation is prone to human error. Artificial Intelligence-powered Internet of Things solutions can reduce human intervention to a certain extent, realizing more scalable environments to handle the food production problem. In this research, an attention-based Temporal Fusion Transformers deep learning model was proposed and validated to forecast nitrate levels in an aquaponics environment. An aquaponics dataset with temporal features and a high number of input lines has been employed for validation and extensive analysis. Experimental results demonstrate significant improvements of the proposed model over baseline models in terms of MAE, MSE, and Explained Variance metrics considering one-hour sequences. Utilizing the proposed solution can help enhance the automation of aquaponics environments.
Ahmet Metin, Ahmet Kasif, Cagatay Catal
J. Supercomput.3
2022 Product failure detection for production lines using a data-driven model
Ziqiu Kang, Cagatay Catal, Bedir Tekinerdogan
Expert Syst. Appl.2
2022 Predictive maintenance using digital twins: A systematic literature review
abstract
Predictive maintenance is a technique for creating a more sustainable, safe, and profitable industry. One of the key challenges for creating predictive maintenance systems is the lack of failure data, as the machine is frequently repaired before failure. Digital Twins provide a real-time representation of the physical machine and generate data, such as asset degradation, which the predictive maintenance algorithm can use. Since 2018, scientific literature on the utilization of Digital Twins for predictive maintenance has accelerated, indicating the need for a thorough review. This research aims to gather and synthesize the studies that focus on predictive maintenance using Digital Twins to pave the way for further research. A systematic literature review (SLR) using an active learning tool is conducted on published primary studies on predictive maintenance using Digital Twins, in which 42 primary studies have been analyzed. This SLR identifies several aspects of predictive maintenance using Digital Twins, including the objectives, application domains, Digital Twin platforms, Digital Twin representation types, approaches, abstraction levels, design patterns, communication protocols, twinning parameters, and challenges and solution directions. These results contribute to a Software Engineering approach for developing predictive maintenance using Digital Twins in academics and the industry. This study is the first SLR in predictive maintenance using Digital Twins. We answer key questions for designing a successful predictive maintenance model leveraging Digital Twins. We found that to this day, computational burden, data variety, and complexity of models, assets, or components are the key challenges in designing these models.
Raymon van Dinter, Bedir Tekinerdogan, Cagatay Catal
Inf. Softw. Technol.3
2022 Obstacles of On-Premise Enterprise Resource Planning Systems and Solution Directions
abstract
The article presents the results of a Systematic Literature Review (SLR) that has been carried out to identify and present the state-of-the-art of ERP systems, describe the obstacles of on-premise ERP systems, and provide general solutions to tackle these challenges. Based on this SLR, 22 obstacles are identified, the dependencies and interactions among these obstacles are described, and finally the corresponding solutions as described in the primary studies are discussed in detail. Our study shows that there is a general agreement on the obstacles of on-premise ERP systems and further research is needed to provide satisfactory solutions to the obstacles.
Senem Sancar Gozukara, Bedir Tekinerdogan, Cagatay Catal
J. Comput. Inf. Syst.3
2022 Applications of deep learning for phishing detection: a systematic literature review
Cagatay Catal, Görkem Giray, Bedir Tekinerdogan, Sandeep Kumar 0004, Suyash Shukla
Knowl. Inf. Syst.1
2022 Applications of deep learning for mobile malware detection: A systematic literature review
Cagatay Catal, Görkem Giray, Bedir Tekinerdogan
Neural Comput. Appl.1
2022 A Firewall Policy Anomaly Detection Framework for Reliable Network Security
abstract
One of the key challenges in computer networks is network security. For securing the network, various solutions have been proposed, including network security protocols and firewalls. In the case of so-called packet-filtering firewalls, policy rules are implemented to monitor changes to the network and preserve the required security level. Due to the dramatic increase of devices, however, and herewith the rapid increase of the size of the policy rules, firewall policy anomalies occur more frequently. This requires careful implementation of the policy rules to ensure cost-efficient solutions for anomaly detection to support network security. In this study, we present an anomaly detection framework for detecting intrafirewall policy anomaly rules. The framework supports the simulation of packets through the firewall ruleset for validating and enhancing the security level of the network. The framework is validated using four different types of firewall policy anomalies. Experimental results demonstrate that the framework is effective and efficient in detecting firewall policy anomalies.
Cengiz Togay, Ahmet Kasif, Cagatay Catal, Bedir Tekinerdogan
IEEE Trans. Reliab.3
2021 A feature-based approach for guiding the selection of Internet of Things cybersecurity standards using text mining
abstract
Abstract Cybersecurity is critical in realizing Internet of Things (IoT) applications and many different standards have been introduced specifically for this purpose. However, selecting relevant standards is not trivial and requires a broad understanding of cybersecurity and knowledge about the available standards. In this study, we present a systematic approach that guides IoT system developers in selecting relevant cybersecurity standards for their IoT projects. The systematic approach has been developed in four stages. First, the common and variant features of IoT cybersecurity have been modeled using a feature model. Second, an up‐to‐date overview of the IoT cybersecurity standards landscape has been mapped by combining existing overviews. Third, a text mining algorithm has been implemented. Fourth, the systematic approach has been modeled using business process modeling notation. Our case study demonstrated that this approach is effective and efficient for guiding the selection of IoT cybersecurity standards.
Koen van der Schaaf, Bedir Tekinerdogan, Cagatay Catal
Concurr. Comput. Pract. Exp.3
2021 A decision support system for automating document retrieval and citation screening
abstract
The systematic literature review (SLR) process includes several steps to collect secondary data and analyze it to answer research questions. In this context, the document retrieval and primary study selection steps are heavily intertwined and known for their repetitiveness, high human workload, and difficulty identifying all relevant literature. This study aims to reduce human workload and error of the document retrieval and primary study selection processes using a decision support system (DSS). An open-source DSS is proposed that supports the document retrieval step, dataset preprocessing, and citation classification. The DSS is domain-independent, as it has proven to carefully select an article’s relevance based solely on the title and abstract. These features can be consistently retrieved from scientific database APIs. Additionally, the DSS is designed to run in the cloud without any required programming knowledge for reviewers. A Multi-Channel CNN architecture is implemented to support the citation screening process. With the provided DSS, reviewers can fill in their search strategy and manually label only a subset of the citations. The remaining unlabeled citations are automatically classified and sorted based on probability. It was shown that for four out of five review datasets, the DSS's use achieved significant workload savings of at least 10%. The cross-validation results show that the system provides consistent results up to 88.3% of work saved during citation screening. In two cases, our model yielded a better performance over the benchmark review datasets. As such, the proposed approach can assist the development of systematic literature reviews independent of the domain. The proposed DSS is effective and can substantially decrease the document retrieval and citation screening steps' workload and error rate.
Raymon van Dinter, Cagatay Catal, Bedir Tekinerdogan
Expert Syst. Appl.2
2021 Automation of systematic literature reviews: A systematic literature review
Raymon van Dinter, Bedir Tekinerdogan, Cagatay Catal
Inf. Softw. Technol.3
2019 Native Code Generation as a Service
abstract
With the widespread use of mobile applications in daily life, it has become crucial for enterprise software companies to quickly develop these applications for multiple platforms. Cross-platform mobile application development is one of the most adopted solutions for rapid development. Since most of these solutions do not generate native code for the underlying platform, the artefacts generally do not satisfy the requirements defined at the beginning of the project. This study designed and implemented a native code generation framework called Nativator built as a cloud service. The framework, which is capable of producing native code for iOS and Android platforms using web-based user interfaces, was implemented based on an open source compiler platform called “Roslyn”. Four case studies were performed to analyze the execution performance of the applications built with the proposed framework. The experimental results demonstrated that the execution performance of the applications built with Nativator is comparable with the applications generated via the state-of-the-art mobile application development framework called Xamarin. Because this framework was implemented as a cloud service, it has several advantages over traditional approaches such as access from anywhere, no installation and flexible and more resources from cloud infrastructure.
Akhan Akbulut, Cagatay Catal, Emre Karadeniz, Emre Turgut
Int. J. Softw. Eng. Knowl. Eng.2
2019 Aligning software engineering education with industrial needs: A meta-analysis
Vahid Garousi, Görkem Giray, Eray Tüzün, Cagatay Catal, Michael Felderer
J. Syst. Softw.4
2019 The impact of feature types, classifiers, and data balancing techniques on software vulnerability prediction models
abstract
Abstract Software vulnerabilities form an increasing security risk for software systems, that might be exploited to attack and harm the system. Some of the security vulnerabilities can be detected by static analysis tools and penetration testing, but usually, these suffer from relatively high false positive rates. Software vulnerability prediction (SVP) models can be used to categorize software components into vulnerable and neutral components before the software testing phase and likewise increase the efficiency and effectiveness of the overall verification process. The performance of a vulnerability prediction model is usually affected by the adopted classification algorithm, the adopted features, and data balancing approaches. In this study, we empirically investigate the effect of these factors on the performance of SVP models. Our experiments consist of four data balancing methods, seven classification algorithms, and three feature types. The experimental results show that data balancing methods are effective for highly unbalanced datasets, text‐based features are more useful, and ensemble‐based classifiers provide mostly better results. For smaller datasets, Random Forest algorithm provides the best performance and for the larger datasets, RusboostTree achieves better performance.
Aydin Kaya 0001, Ali Seydi Keçeli, Cagatay Catal, Bedir Tekinerdogan
J. Softw. Evol. Process.3
2017 Automatic Software Categorization Using Ensemble Methods and Bytecode Analysis
abstract
Software repositories consist of thousands of applications and the manual categorization of these applications into domain categories is very expensive and time-consuming. In this study, we investigate the use of an ensemble of classifiers approach to solve the automatic software categorization problem when the source code is not available. Therefore, we used three data sets (package level/class level/method level) that belong to 745 closed-source Java applications from the Sharejar repository. We applied the Vote algorithm, AdaBoost, and Bagging ensemble methods and the base classifiers were Support Vector Machines, Naive Bayes, J48, IBk, and Random Forests. The best performance was achieved when the Vote algorithm was used. The base classifiers of the Vote algorithm were AdaBoost with J48, AdaBoost with Random Forest, and Random Forest algorithms. We showed that the Vote approach with method attributes provides the best performance for automatic software categorization; these results demonstrate that the proposed approach can effectively categorize applications into domain categories in the absence of source code.
Cagatay Catal, Serkan Tugul, Basar Akpinar
Int. J. Softw. Eng. Knowl. Eng.1
2013 Test case prioritization: a systematic mapping study
Cagatay Catal, Deepti Mishra 0001
Softw. Qual. J.1
2011 A Composite Project Effort Estimation Approach in an Enterprise Software Development Project
Cagatay Catal, Mehmet S. Aktas
SEKE1
2011 Thresholds based outlier detection approach for mining class outliers: An empirical case study on software measurement datasets
Oral Alan, Cagatay Catal
Expert Syst. Appl.2
2011 Software fault prediction: A literature review and current trends
Cagatay Catal
Expert Syst. Appl.1
2011 Practical development of an Eclipse-based software fault prediction tool using Naive Bayes algorithm
Cagatay Catal, Ugur Sevim, Banu Diri
Expert Syst. Appl.1
2011 Class noise detection based on software metrics and ROC curves
Cagatay Catal, Oral Alan, Kerime Balkan
Inf. Sci.1
2009 Unlabelled extra data do not always mean extra performance for semi-supervised fault prediction
abstract
Abstract: This research focused on investigating and benchmarking several high performance classifiers called J48, random forests, naive Bayes, KStar and artificial immune recognition systems for software fault prediction with limited fault data. We also studied a recent semi‐supervised classification algorithm called YATSI (Yet Another Two Stage Idea) and each classifier has been used in the first stage of YATSI. YATSI is a meta algorithm which allows different classifiers to be applied in the first stage. Furthermore, we proposed a semi‐supervised classification algorithm which applies the artificial immune systems paradigm. Experimental results showed that YATSI does not always improve the performance of naive Bayes when unlabelled data are used together with labelled data. According to experiments we performed, the naive Bayes algorithm is the best choice to build a semi‐supervised fault prediction model for small data sets and YATSI may improve the performance of naive Bayes for large data sets. In addition, the YATSI algorithm improved the performance of all the classifiers except naive Bayes on all the data sets.
Cagatay Catal, Banu Diri
Expert Syst. J. Knowl. Eng.1
2009 A systematic review of software fault prediction studies
Cagatay Catal, Banu Diri
Expert Syst. Appl.1
2009 Investigating the effect of dataset size, metrics sets, and feature selection techniques on software fault prediction problem
Cagatay Catal, Banu Diri
Inf. Sci.1
2008 A Fault Prediction Model with Limited Fault Data to Improve Test Process
Cagatay Catal, Banu Diri
PROFES1
2008 A Conceptual Framework to Integrate Fault Prediction Sub-Process for Software Product Lines
abstract
Software product line engineering is a growing recent paradigm to develop similar products using reusable core assets such as architecture and test cases. The general aim is to enhance quality and decrease development costs. Current software product line engineering frameworks apply only a few quality assurance activities but today's single-system engineering has much more quality assurance activities that can be adapted to software product lines. In this study, software fault prediction sub- process is integrated into software product line engineering framework and the key activities are defined. This approach will improve quality and enhance testing process for software product lines.
Cagatay Catal, Banu Diri
TASE1
2007 Software Fault Prediction with Object-Oriented Metrics Based Artificial Immune Recognition System
Cagatay Catal, Banu Diri
PROFES1