VLDB 2026 Research / reviewers in the wild / expert
Miroslaw Ochodek
dblp:39/2146
· DBLP profile ↗
28ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-9103-717XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging LLM-based data augmentation for automatic classification of recurring tasks in software development projects
Wlodzimierz Wysocki, Miroslaw Ochodek |
J. Syst. Softw. | 2 |
| 2025 | Design pattern recognition: a study of large language modelsabstractAbstract Context As Software Engineering (SE) practices evolve due to extensive increases in software size and complexity, the importance of tools to analyze and understand source code grows significantly. Objective This study aims to evaluate the abilities of Large Language Models (LLMs) in identifying DPs in source code, which can facilitate the development of better Design Pattern Recognition (DPR) tools. We compare the effectiveness of different LLMs in capturing semantic information relevant to the DPR task. Methods We studied Gang of Four (GoF) DPs from the P-MARt repository of curated Java projects. State-of-the-art language models, including Code2Vec, CodeBERT, CodeGPT, CodeT5, and RoBERTa, are used to generate embeddings from source code. These embeddings are then used for DPR via a k-nearest neighbors prediction. Precision, recall, and F1-score metrics are computed to evaluate performance. Results RoBERTa is the top performer, followed by CodeGPT and CodeBERT, which showed mean F1 Scores of 0.91, 0.79, and 0.77, respectively. The results show that LLMs without explicit pre-training can effectively store semantics and syntactic information, which can be used in building better DPR tools. Conclusion The performance of LLMs in DPR is comparable to existing state-of-the-art methods but with less effort in identifying pattern-specific rules and pre-training. Factors influencing prediction performance in Java files/programs are analyzed. These findings can advance software engineering practices and show the importance and abilities of LLMs for effective DPR in source code. Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron, Miroslaw Ochodek, Darko Durisic |
Empir. Softw. Eng. | 5 |
| 2024 | Automatic Classification of Recurring Tasks in Software Development ProjectsabstractBackground: Information about project tasks stored in Issue tracking systems (ITS) can be used for project analytics or process simulation. Such tasks can be categorized as stateful or recurrent. Although automatic categorization of stateful tasks is relatively simple, performing the same task for recurrent tasks constitutes a challenge. Aims: The goal of this study is to investigate the possibility of employing machine-learning algorithms to automatically categorize recurrent tasks in software projects based on information stored in ITS. Method: We perform a study on a dataset from six industrial projects containing 9,589 tasks and augment it with an additional dataset of 91,145 task descriptions from other industrial projects to up-sample minority classes during training. We perform ten runs of 10-fold cross-validation for each project and evaluate classifiers using a set of state-of-the-art prediction quality metrics, i.e., Accuracy, Precision, Recall, F1-score, and MCC. Our machine-learning pipeline includes a Transformer-based sentence embed-der (‘mxbai-embed-large-vl’) and XGBoost classifier. Results: The model automatically classifies software process tasks into 14 classes with MCC ranging between 0.69 and 0.88 (mean: 0.77). We observed higher prediction quality for the largest projects in the dataset and those managed according to “traditional” project management methodologies. Conclusions: We conclude that machine-learning algorithms can effectively categorize re-current tasks. However, it requires collecting a large balanced dataset of ITS tasks or using a pre-trained model like the one provided in this study. Wlodzimierz Wysocki, Miroslaw Ochodek |
SEAA | 2 |
| 2023 | TransDPR: Design Pattern Recognition Using Programming Language ModelsabstractCurrent Design Pattern Recognition (DPR) methods have limitations, such as the reliance on semantic information, limited recognition of novel or modified pattern versions, and other factors. We present an introductory DPR technique by using a Programming Language Model (PLM) called TransDPR, which utilizes a Facebook pre-trained model (TransCoder), which is a Cross-lingual programming Language Model (XLM) based on a transformer architecture. We leverage an n-dimensional vector representation of programs and apply logistic regression to learn design patterns (DPs). Our approach utilizes the GitHub repository to collect singleton and prototype DP programs written in$C$++ source code. Our results indicate that TransDPR achieves 90% accuracy and an F1-score of 0.88 on open-source projects. We evaluate the proposed model on two developed modules from Volvo Cars and invite the original developers to validate the prediction results. Sushant Kumar Pandey, Miroslaw Staron, Jennifer Horkoff, Miroslaw Ochodek, Nicholas Mucci, Darko Durisic |
ESEM | 4 |
| 2023 | Defect Backlog Size Prediction for Open-Source Projects with the Autoregressive Moving Average and Exponential Smoothing ModelsabstractContext: predicting the number of defects in a defect backlog in a given time horizon can help allocate project resources and organize software development.Goal: to compare the accuracy of three defect backlog prediction methods in the context of large open-source (OSS) projects, i.e., ARIMA, Exponential Smoothing (ETS), and the state-of-the-art method developed at Ericsson AB (MS).Method: we perform a simulation study on a sample of 20 open-source projects to compare the prediction accuracy of the methods.Also, we use the Naïve prediction method as a baseline for sanity check.We use statistical inference tests and effect size coefficients to compare the prediction errors.Results: ARIMA, ETS, and MS were more accurate than the Naïve method.Also, the prediction errors were statistically lower for ETS than for MS (however, the effect size was negligible).Conclusions: ETS seems slightly more accurate than MS when predicting defect backlog size of OSS projects. Paulina Aniola, Sushant Kumar Pandey, Miroslaw Staron, Miroslaw Ochodek |
FedCSIS | 4 |
| 2023 | On the Applicability of the Pareto Principle to Source-Code Growth in Open Source ProjectsabstractContext: research on understanding the laws related to software-project evolution can indirectly impact the way we design software development processes, e.g., knowing the nature of the code-repository content growth could help us improve the ways we monitor the progress of OSS software development projects and predict their future development Goal: our aim is to empirically verify a hypothesis that the OSS code repositories grow in size according to the Pareto principle.Method: we collected and curated a sample of 31,343 OSS code repositories hosted on GitHub and analyzed their content growth over time to verify whether it follows the Pareto principle.Results: we observed that, on average, monotonically growing OSS repositories reach 75% of their final content size within the first 25% revisions.Conclusions: the content size of monotonically growing OSS repositories seems to grow in size according to the Pareto principle with the 75/25 ratio. Korneliusz Szymanski, Miroslaw Ochodek |
FedCSIS | 2 |
| 2023 | Towards reliable rule mining about code smells: The McPython approach (Invited Lecture - Extended Abstract)
Maciej Ziobrowski, Miroslaw Ochodek, Jerzy R. Nawrocki, Bartosz Walter |
FedCSIS | 2 |
| 2022 | On Testing Security Requirements in Industry - A Survey Study
Sylwia Kopczynska, Daniel Craviee De Abreu Vieira, Miroslaw Ochodek |
REFSQ | 3 |
| 2022 | On the benefits and problems related to using Definition of Done - A survey study
Sylwia Kopczynska, Miroslaw Ochodek, Jakub Piechowiak, Jerzy R. Nawrocki |
J. Syst. Softw. | 2 |
| 2020 | Using Machine Learning to Identify Code Fragments for Manual ReviewabstractCode reviews are one of the first quality assurance tasks in continuous software integration and delivery. The goal of our work is to reduce the need for manual reviews by automatically identify which code fragments should be further reviewed manually. We conducted an action research study with two companies where we extracted code reviews and build machine learning classifiers (AdaBoost and Convolutional Neural Network– CNN). Our results show that the accuracy of recognizing code fragments that require manual review, measured with Matthews Correlation Coefficient, was 0.70 in the combination of our own feature extraction and CNN. We conclude that this way of combining automation with manual code reviews can improve the speed of reviews while providing organizations with the possibility to support knowledge transfer among the designers. Miroslaw Staron, Miroslaw Ochodek, Wilhelm Meding, Ola Soder |
SEAA | 2 |
| 2020 | A Case Study on a Hybrid Approach to Assessing the Maturity of Requirements Engineering Practices in Agile Projects (REMMA)
Miroslaw Ochodek, Sylwia Kopczynska, Jerzy R. Nawrocki |
SOFSEM | 1 |
| 2020 | Maintainability of Automatic Acceptance Tests for Web Applications - A Case Study Comparing Two Approaches to Organizing Code of Test Cases
Aleksander Sadaj, Miroslaw Ochodek, Sylwia Kopczynska, Jerzy R. Nawrocki |
SOFSEM | 2 |
| 2020 | Recognizing lines of code violating company-specific coding guidelines using machine learningabstractAbstract Software developers in big and medium-size companies are working with millions of lines of code in their codebases. Assuring the quality of this code has shifted from simple defect management to proactive assurance of internal code quality. Although static code analysis and code reviews have been at the forefront of research and practice in this area, code reviews are still an effort-intensive and interpretation-prone activity. The aim of this research is to support code reviews by automatically recognizing company-specific code guidelines violations in large-scale, industrial source code. In our action research project, we constructed a machine-learning-based tool for code analysis where software developers and architects in big and medium-sized companies can use a few examples of source code lines violating code/design guidelines (up to 700 lines of code) to train decision-tree classifiers to find similar violations in their codebases (up to 3 million lines of code). Our action research project consisted of (i) understanding the challenges of two large software development companies, (ii) applying the machine-learning-based tool to detect violations of Sun’s and Google’s coding conventions in the code of three large open source projects implemented in Java, (iii) evaluating the tool on evolving industrial codebase, and (iv) finding the best learning strategies to reduce the cost of training the classifiers. We were able to achieve the average accuracy of over 99% and the average F-score of 0.80 for open source projects when using ca. 40K lines for training the tool. We obtained a similar average F-score of 0.78 for the industrial code but this time using only up to 700 lines of code as a training dataset. Finally, we observed the tool performed visibly better for the rules requiring to understand a single line of code or the context of a few lines (often allowing to reach the F-score of 0.90 or higher). Based on these results, we could observe that this approach can provide modern software development companies with the ability to use examples to teach an algorithm to recognize violations of code/design guidelines and thus increase the number of reviews conducted before the product release. This, in turn, leads to the increased quality of the final software. Miroslaw Ochodek, Regina Hebig, Wilhelm Meding, Gert Frost, Miroslaw Staron |
Empir. Softw. Eng. | 1 |
| 2020 | PHANTOM: Curating GitHub for engineered software projects using time-series clusteringabstractAbstract Context Within the field of Mining Software Repositories, there are numerous methods employed to filter datasets in order to avoid analysing low-quality projects. Unfortunately, the existing filtering methods have not kept up with the growth of existing data sources, such as GitHub, and researchers often rely on quick and dirty techniques to curate datasets. Objective The objective of this study is to develop a method capable of filtering large quantities of software projects in a resource-efficient way. Method This study follows the Design Science Research (DSR) methodology. The proposed method, PHANTOM, extracts five measures from Git logs. Each measure is transformed into a time-series, which is represented as a feature vector for clustering using the k-means algorithm. Results Using the ground truth from a previous study, PHANTOM was shown to be able to rediscover the ground truth on the training dataset, and was able to identify “engineered” projects with up to 0.87 Precision and 0.94 Recall on the validation dataset. PHANTOM downloaded and processed the metadata of 1,786,601 GitHub repositories in 21.5 days using a single personal computer, which is over 33% faster than the previous study which used a computer cluster of 200 nodes. The possibility of applying the method outside of the open-source community was investigated by curating 100 repositories owned by two companies. Conclusions It is possible to use an unsupervised approach to identify engineered projects. PHANTOM was shown to be competitive compared to the existing supervised approaches while reducing the hardware requirements by two orders of magnitude. Peter Pickerill, Heiko Joshua Jungen, Miroslaw Ochodek, Michal Mackowiak, Miroslaw Staron |
Empir. Softw. Eng. | 3 |
| 2020 | Deep learning model for end-to-end approximation of COSMIC functional size based on use-case names
Miroslaw Ochodek, Sylwia Kopczynska, Miroslaw Staron |
Inf. Softw. Technol. | 1 |
| 2019 | When NFR Templates Pay Back? A Study on Evolution of Catalog of NFR Templates
Sylwia Kopczynska, Jerzy R. Nawrocki, Miroslaw Ochodek |
PROFES | 3 |
| 2019 | Simsax: A measure of project similarity based on symbolic approximation method and software defect inflow
Miroslaw Ochodek, Miroslaw Staron, Wilhelm Meding |
Inf. Softw. Technol. | 1 |
| 2018 | An empirical study on catalog of non-functional requirement templates: Usefulness and maintenance issues
Sylwia Kopczynska, Jerzy R. Nawrocki, Miroslaw Ochodek |
Inf. Softw. Technol. | 3 |
| 2018 | On some end-user programming constructs and their understandability
Michal Mackowiak, Jerzy R. Nawrocki, Miroslaw Ochodek |
J. Syst. Softw. | 3 |
| 2018 | Perceived importance of agile requirements engineering practices - A survey
Miroslaw Ochodek, Sylwia Kopczynska |
J. Syst. Softw. | 1 |
| 2016 | Towards Semi-Automatic Size Measurement of User Interfaces in Web Applications with IFPUG SNAPabstractSoftware Non-functional Assessment Process is a non-functional size measurement method proposed by the International Function Point Users Group. It can be used to measure the size of non-functional requirements related to usability of graphical user interfaces (UI). Unfortunately, measuring such requirements seems time-consuming because it requires identifying all UI elements and graphical properties that were configured to meet these requirements. In this paper, we propose a semi-automated approach to measure size of non-functional related to user interfaces of web applications. The method takes as an input a set of exemplary screens of application (HTML and CSS) and rules describing the mapping between HTML elements and UI Elements. The method provides a list of UI Elements and graphical properties that were configured as an output. We preliminarily validated the proposed method using a prototype tool that we had developed. Hassan Mansoor, Miroslaw Ochodek |
IWSM-Mensura | 2 |
| 2016 | Approximation of COSMIC Functional Size of Scenario-Based Requirements in Agile Based on Syntactic Linguistic Features - A Replication StudyabstractContext: Expert judgment is the most frequently used method of effort estimation in Agile software development. Unfortunately, Agile teams often underestimate development effort. Therefore, it seems beneficial to support such teams with the information regarding the functional size of requirements they are estimating. Hussain, Kosseim and Ormandjieva (HKO) proposed a method that can be used to automatically classify textual requirements with respect to their COSMIC functional size. Unfortunately, the method has not been sufficiently validated to confirm its usefulness. Objective: To provide external validation of the HKO method and investigate if it can be applied to classify scenario-based requirements (in the form of use cases) with respect to their COSMIC size. Method: Similarily to the original study, we used a set of natural language processing tools to extract syntactic linguistic features and the C4.5 decision tree-based classifiers to classify requirements. We validated the performance of the classifiers using the 10-fold cross-validation procedure on a dataset containing 93 use cases. We compared the performance of the HKO method with the performance of the classifiers trained using a single prediction feature-the number of steps in a use case. Results: Depending on the considered number of size classes and the algorithm used to compute boundaries of the classes, the accuracy of the HKO method ranged between .387 and .785 while the Cohen's kappa index was between .194 and .577. The accuracy of the use-case-steps-based classifiers performed slightly worse. Their accuracy ranged between .015 and .769 while Cohen's kappa was between .067 and .423. We observed that the performance of both types of classifiers dropped visibly when applied to four or more size classes. Conclusion: The classification performance of the HKO method was moderate. However, it was still better than the classification based on the number of steps. Unfortunately, we also observed that the accuracy of the HKO method is sensitive to the language used in descriptions of requirements. Miroslaw Ochodek |
IWSM-Mensura | 1 |
| 2016 | Functional size approximation based on use-case names
Miroslaw Ochodek |
Inf. Softw. Technol. | 1 |
| 2015 | HAZOP-based identification of events in use casesabstractAbstract Completeness is one of the main quality attributes of requirements specifications. If functional requirements are expressed as use cases, one can be interested in event completeness. A use case is event complete if it contains description of all the events that can happen when executing the use case. Missing events in any use case can lead to higher project costs. Thus, the question arises of what is a good method of identification of events in use cases and what accuracy and review speed one can expect from it. The goal of this study was to check if (1) HAZOP-based event identification is more effective than ad hoc review and (2) what is the review speed of these two approaches. Two controlled experiments were conducted in order to evaluate ad hoc approach and H4U method to event identification. The first experiment included 18 students, while the second experiment was conducted with the help of 82 professionals. In both cases, accuracy and review speed of the investigated methods were measured and analyzed. Moreover, the usage of HAZOP keywords was analyzed. In both experiments, a benchmark specification based on use cases was used. The first experiment with students showed that a HAZOP-based review is more effective in event identification than ad hoc review and this result is statistically significant. However, the reviewing speed of HAZOP-based reviews is lower. The second experiment with professionals confirmed these results. These experiments showed also that event completeness is hard to achieve. It on average ranged from 0.15 to 0.26. HAZOP-based identification of events in use cases is an useful alternative to ad hoc reviews. It can achieve higher event completeness at the cost of an increase in effort. Jakub Jurkiewicz, Jerzy R. Nawrocki, Miroslaw Ochodek, Tomasz Glowacki |
Empir. Softw. Eng. | 3 |
| 2014 | Agile Requirements Engineering: A Research Perspective
Jerzy R. Nawrocki, Miroslaw Ochodek, Jakub Jurkiewicz, Sylwia Kopczynska, Bartosz Alchimowicz |
SOFSEM | 2 |
| 2011 | Improving the reliability of transaction identification in use cases
Miroslaw Ochodek, Bartosz Alchimowicz, Jakub Jurkiewicz, Jerzy R. Nawrocki |
Inf. Softw. Technol. | 1 |
| 2011 | Simplifying effort estimation based on Use Case Points
Miroslaw Ochodek, Jerzy R. Nawrocki, K. Kwarciak |
Inf. Softw. Technol. | 1 |
| 2008 | 3-step knowledge transition: a case study on architecture evaluationabstractSoftware Engineering is developing very fast. To keep up with the changes, software companies need effective methods of knowledge transfer. In the paper a 3-step approach to knowledge transfer, called Technical Drama, is presented. The paper is focused on transferring knowledge concerning architecture evaluation, but the approach could also be applied to transferring knowledge concerning inspections, testing etc. It is claimed in the paper that the Technical Drama can be useful in the industrial context (two case studies are described) as well as at university (then a kind of software studio is required). Bartosz Michalik, Jerzy R. Nawrocki, Miroslaw Ochodek |
ICSE | 3 |