VLDB 2026 Research / reviewers in the wild / expert
Önder Babur
dblp:165/0091
· DBLP profile ↗
18ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-1460-2825ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing research software engineering with AI: a research frameworkabstractAbstract The rapid adoption of Artificial Intelligence (AI) and Generative AI (GenAI) tools is transforming the creation, maintenance, and dissemination of research software. Despite their growing prevalence, the implications of these technologies for Research Software Engineering (RSE) practices remain underexplored. This work introduces AI4RSE , an emerging research domain focused on the integration of AI into the development lifecycle of research software. To investigate current trends in AI-augmented RSE, we conducted an empirical study of more than 1,500 open-source research software repositories hosted on Zenodo. Each repository was assessed using a quadrant-based typology defined by two key dimensions: software engineering maturity and the level of AI integration. Our analysis combined static and semantic code inspection, evaluation of alignment with the FAIR Principles for Research Software (FAIR4RS), and heuristic classification of generative AI usage and MLOps adoption. Repositories are categorized into four development modes: Exploratory Coding , Vibe Coding , RSE , and AI4RSE , which reflect different levels of process rigor and AI tool integration. While many projects exhibit informal development patterns, a growing subset demonstrates mature, AI-assisted workflows. This landscape reveals key challenges, such as reproducibility risks and licensing ambiguity, while also highlighting emerging opportunities, including AI-assisted testing and intelligent documentation generation. The findings support a research agenda for AI4RSE, outlining benchmarks, guidelines, and community standards to promote responsible, reproducible, and scalable adoption of AI in scientific software development. Siamak Farshidi, Kwabena Ebo Bennin, Önder Babur, June Sallou, Ayalew Kassahun, Bedir Tekinerdogan |
Autom. Softw. Eng. | 3 |
| 2025 | An empirical study of business process models and model clones on GitHubabstractBusiness process management entails a multi-billion-dollar industry that is founded on modeling business processes to analyze, understand, improve, and automate them. Business processes consist of a set of interconnected activities that an organization follows to achieve its goals and objectives. While the existence of business process models in open source has been reported in the literature, there is little work in characterizing their landscape. This paper presents the first characterization of business process models in open source, particularly on GitHub. The landscape is formed by 25,866 business process models across 4,954 repositories, with 16% of the repositories belonging to organizations. We discover that models belong to at least 16 domains including traditional software, machine learning, sales, business services, and financial services. These models are created using at least 28 different tools. Our exploration into cloning among the models shows that about 90% of all models are clones of each other. Application domains such as machine learning, traditional software, and business services demonstrate a higher occurrence of clones while in another dimension, clones are found across more repositories owned by industry as compared to those owned by academia. Also, contrary to code clones, we find that the majority of process model cloning occurs across multiple repositories. While our study acts as a precursor for future efforts to develop effective modeling practices in the field of business processes, it also emphasizes the need to address cloning and its implications in the context of reuse, maintenance, and modeling approaches. Mahdi Saeedi Nikoo, Sangeeth Kochanthara, Önder Babur, Mark van den Brand |
Empir. Softw. Eng. | 3 |
| 2025 | Domain-Driven Design in software development: A systematic literature review on implementation, challenges, and effectivenessabstractContext: Domain-Driven Design (DDD) has gained significant attention in software development for its potential to address complex software challenges, particularly in the areas of system refactoring, reimplementation, and adoption. Using domain knowledge, DDD aims to solve complex business problems effectively. Objective: This SLR aims to provide an analysis of existing research on DDD in software development, paint a picture of DDD in solving software problems, identify the challenges encountered during its application and explore the results of these studies. Method: We systematically selected 36 peer reviewed studies and conducted quantitative and qualitative analyzes to synthesize the findings. Results: DDD has effectively improved software systems, with its key concepts. The application of DDD in microservices has gained prominence for its ability to facilitate system decomposition. Some studies lacked empirical evaluations, highlighting challenges in onboarding and the need for expertise. Conclusion: Adopting DDD benefits software development, involving stakeholders such as engineers, architects, managers, and domain experts. More empirical evaluations and open discussions on challenges are needed. Collaboration between academia and industry advances the adoption and transfer of knowledge of DDD in projects. Ozan Özkan, Önder Babur, Mark van den Brand |
J. Syst. Softw. | 2 |
| 2024 | Language usage analysis for EMF metamodels on GitHubabstractAbstract Context EMF metamodels lie at the heart of model-based approaches for a variety of tasks, notably for defining the abstract syntax of modeling languages. The language design of EMF metamodels itself is part of a design process, where the needs of its specific range of users should be satisfied. Studying how people actually use the language in the wild would enable empirical feedback for improving the design of the EMF metamodeling language. Objective Our goal is to study the language usage of EMF metamodels in public engineered projects on GitHub. We aim to reveal information about the usage of specific language constructs, whether they match the language design. Based on our findings, we plan to suggest improvements in the EMF metamodelling language. Method We adopt a sample study research strategy and collect data from the EMF metamodels on GitHub. After a series of preprocessing steps including filtering out non-engineered projects and deduplication, we employ an analytics workflow on top of a graph database to formulate generalizing statements about the artifacts under study. Based on the results, we also give actionable suggestions for the EMF metamodeling language design. Results We have conducted various analyses on metaclass, attribute, feature/relationship usage as well as specific parts of the language: annotations and generics. Our findings reveal that the most used metaclasses are not the main building blocks of the language, but rather auxiliary ones. Some of the metaclasses, metaclass features and relations are almost never used. There are a few attributes which are almost exclusively used with a single value or illegal values. Some of the language features such as special forms of generics are very rarely used. Based on our findings, we provide suggestions to improve the EMF language, e.g. removing a language element, restricting its values or refining the metaclass hierarchy. Conclusions In this paper, we present an extensive empirical study into the language usage of EMF metamodels on GitHub. We believe this study fills a gap in the literature of model analytics and will hopefully help future improvement of the EMF metamodeling language. Önder Babur, Eleni Constantinou, Alexander Serebrenik |
Empir. Softw. Eng. | 1 |
| 2024 | A systematic review on food recommender systemsabstractThe Internet has revolutionized the way information is retrieved, and the increase in the number of users has resulted in a surge in the volume and heterogeneity of available data. Recommender systems have become popular tools to help users retrieve relevant information quickly. Food Recommender Systems (FRS), in particular, have proven useful in overcoming the overload of information present in the food domain. However, the recommendation of food is a complex domain with specific characteristics causing many challenges. Additionally, very few systematic literature reviews have been conducted in the domain on FRS. This paper presents a systematic literature review that summarizes the current state-of-the-art in FRS. Our systematic review examines the different methods and algorithms used for recommendation, the data and how it is processed, and evaluation methods. It also presents the advantages and disadvantages of FRS. To achieve this, a total of 67 high-quality studies were selected from a pool of 2,738 studies using strict quality criteria. The review provides valuable information to the research field, helping researchers in the domain to select a strategy to develop FRS. This review can help improve the efficiency of development, thus closing the gap between the development of FRS and other recommender systems. Jon Nicolas Bondevik, Kwabena Ebo Bennin, Önder Babur, Carsten Ersch |
Expert Syst. Appl. | 3 |
| 2024 | Generating domain models from natural language text using NLP: a benchmark dataset and experimental comparison of tools
Fatma Bozyigit, Tolgahan Bardakci, Alireza Khalilipour, Moharram Challenger, Guus Ramackers, Önder Babur, Michel R. V. Chaudron |
Softw. Syst. Model. | 6 |
| 2023 | Refactoring with domain-driven design in an industrial contextabstractAbstract Context Software developers need to constantly work on evolving the structure and the stability of the code due to changing business needs of the product. There are various refactoring approaches in industry which promise improvements over source code composition and maintainability. Objective In our research, we want to improve the maintainability of an existing system through refactoring using Domain-Driven Design (DDD) as a software design approach. We also aim for providing empirical evidence on its effect on maintainability and the challenges as perceived by developers. Method In this study, we applied the action research methodology, which facilitates close academia-industry collaboration and regular presence in the studied product. We utilized focus groups to discover problems of the existing system with a qualitative approach. We reviewed the subject codebase to construct our own expert opinion as well and identified problems in the codebase and matched them with the ones raised by engineers in the team. We refactored the existing software system according to DDD principles. To measure the effects of our actions, we utilized Technology Acceptance Model (mTAM) questionnaire, and also semi-structured interviews with the development team for data collection, and card sorting methodology for qualitative analysis. For minimizing bias that might affect our results with the existing software engineers in the team, we extended our measurement with three new joiner software engineers in the team through the think aloud protocol. Results We have identified that engineers mostly gave positive answers to our interview questions, which are mapped to software maintainability metrics defined by ISO/IEC 25010. Our DDD refactoring scored 85 in PU and 83 in PEU, leading to an overall mTAM score of 84. This means acceptable on the acceptability scale, B on the grade scale, and good on the adjective rating scale. Conclusion Our research led us to conclude that a powerful design approach, like DDD, is an effective tool for restructuring and resolving software issues in this situation. It offers standardization to the software and the refactoring efforts. We realized that DDD entails a certain degree of complexity and cognitive load, which is a barrier for software engineers, but they are aware of its benefits. Ozan Özkan, Önder Babur, Mark van den Brand |
Empir. Softw. Eng. | 2 |
| 2023 | On the use of deep learning in software defect predictionabstractAutomated software defect prediction (SDP) methods are increasingly applied, often with the use of machine learning (ML) techniques. Yet, the existing ML-based approaches require manually extracted features, which are cumbersome, time consuming and hardly capture the semantic information reported in bug reporting tools. Deep learning (DL) techniques provide practitioners with the opportunities to automatically extract and learn from more complex and high-dimensional data. The purpose of this study is to systematically identify, analyze, summarize, and synthesize the current state of the utilization of DL algorithms for SDP in the literature. We systematically selected a pool of 102 peer-reviewed studies and then conducted a quantitative and qualitative analysis using the data extracted from these studies. Main highlights include: (1) most studies applied supervised DL; (2) two third of the studies used metrics as an input to DL algorithms; (3) Convolutional Neural Network is the most frequently used DL algorithm. Based on our findings, we propose to (1) develop more comprehensive DL approaches that automatically capture the needed features; (2) use diverse software artifacts other than source code; (3) adopt data augmentation techniques to tackle the class imbalance problem; (4) publish replication packages. Görkem Giray, Kwabena Ebo Bennin, Ömer Köksal, Önder Babur, Bedir Tekinerdogan |
J. Syst. Softw. | 4 |
| 2022 | SAMOS - A framework for model analytics and managementabstractThe increased popularity and adoption of model-* engineering paradigms, such as model-driven and model-based engineering, leads to an increase in the number of models, metamodels, model transformations and other related artifacts. This calls for automated techniques to analyze large collections of those artifacts to manage model-* ecosystems. SAMOS is a framework to address this challenge: it treats model-* artifacts as data, and applies various techniques—ranging from information retrieval to machine learning—to analyze those artifacts in a holistic, scalable and efficient way. Such analyses can help to understand and manage those ecosystems. Önder Babur, Loek Cleophas, Mark van den Brand |
Sci. Comput. Program. | 1 |
| 2020 | DeepClone: Modeling Clones to Generate Code Predictions
Muhammad Hammad 0001, Önder Babur, Hamid Abdul Basit, Mark van den Brand |
ICSR | 2 |
| 2019 | MoCoP: towards a model clone portalabstractWidespread and mature practice of model-driven engineering is leading to a growing number of modeling artifacts and challenges in their management. Model clone detection (MCD) is an important approach for managing and maintaining modeling artifacts. While its counterpart in traditional source code development, code clone detection, is enjoying popularity and more than two decades of development, MCD is still in its infancy in terms of research and tooling. We aim to develop a portal for model clone detection, MoCoP, as a central hub to mitigate adoption barriers and foster MCD research. In this short paper, we present our vision for MoCoP and its features and goals. We discuss MoCoP's key components that we plan on realizing in the short term including public tooling, curated data sets, and a body of MCD knowledge. Our longer term goals include a dedicated service-oriented infrastructure, contests, and forums. We believe MoCoP will strengthen MCD research, tooling, and the community, which in turn will lead to better quality, maintenance, and scalability for model-driven engineering practices. Önder Babur, Matthew Stephan |
MiSE@ICSE | 1 |
| 2018 | Clone Detection for Ecore Metamodels using N-gramsabstractIncreasing model-driven engineering use leads to an abundance of models and metamodels in academic and industrial practice. A key technique for the management and maintenance of those artefacts is model clone detection, where highly similar (meta-)models and (meta-)model fragments are mined from a possibly large amount of data. In this paper we extend the SAMOS framework (Statistical Analysis of MOdelS) to clone detection on Ecore metamodels, using the framework’s n-gram feature extraction, vector space model and clustering capabilities. We perform a case analysis on Ecore metamodels obtained by applying an exhaustive set of single mutations to assess the precision/sensitivity of our technique with respect to various types of mutations. Using mutation analysis, we also briefly evaluate MACH, a comparable UML clone detection tool. Önder Babur |
MODELSWARD | 1 |
| 2018 | Towards Distributed Model Analytics with Apache SparkabstractThe growing number of models and other related artefacts in model-driven engineering has recently led to the emergence of approaches and tools for analyzing and managing them on a large scale. The framework SAMOS applies techniques inspired by information retrieval and data mining to analyze large sets of models. As the data size and analysis complexity goes up, however, further scalability is needed. In this paper we extend SAMOS to operate on Apache Spark, a popular engine for distributed Big Data processing, by partitioning the data and parallelizing the comparison and analysis phase. We present preliminary studies using a cluster infrastructure and report the results for two datasets: one with 250 Ecore metamodels where we detail the performance gain with various settings, and a larger one of 7.3k metamodels with nearly one million model elements for further demonstrating scalability. Önder Babur, Loek Cleophas, Mark van den Brand |
MODELSWARD | 1 |
| 2018 | Improving custom-tailored variability mining using outlier and cluster detection
David Wille, Önder Babur, Loek Cleophas, Christoph Seidl 0001, Mark van den Brand, Ina Schaefer |
Sci. Comput. Program. | 2 |
| 2017 | Using n-grams for the Automated Clustering of Structural Models
Önder Babur, Loek Cleophas |
SOFSEM | 1 |
| 2016 | Hierarchical Clustering of Metamodels for Comparative Analysis and Visualization
Önder Babur, Loek Cleophas, Mark van den Brand |
ECMFA | 1 |
| 2016 | Statistical analysis of large sets of modelsabstractMany applications in Model-Driven Engineering involve processing multiple models, e.g. for comparing and merging of model variants into a common domain model. Despite many sophisticated techniques for model comparison, little attention has been given to the initial data analysis and filtering activities. These are hard to ignore especially in the case of a large dataset, possibly with outliers and sub-groupings. We would like to develop a generic approach for model comparison and analysis for large datasets; using techniques from information retrieval, natural language processing and machine learning. We are implementing our approach as an open framework and have so far evaluated it on public datasets involving domain analysis, repository management and model searching scenarios. Önder Babur |
ASE | 1 |
| 2016 | Towards Statistical Comparison and Analysis of ModelsabstractModel comparison is an important challenge in model-driven engineering, with many application areas such as model versioning and domain model recovery. There are numerous techniques that address this challenge in the literature, ranging from graph-based to linguistic ones. Most of these involve pairwise comparison, which might work, e.g. for model versioning with a small number of models to consider. However, they mostly ignore the case where there is a large number of models to compare, such as in common domain model/metamodel recovery from multiple models. In this paper we present a generic approach for model comparison and analysis as an exploratory first step for model recovery. We propose representing models in vector space model, and applying clustering techniques to compare and analyse a large set of models. We demonstrate our approach on a synthetic dataset of models generated via genetic algorithms. Önder Babur, Loek Cleophas, Tom Verhoeff, Mark van den Brand |
MODELSWARD | 1 |