VLDB 2026 Research / reviewers in the wild / expert
Andriy V. Miranskyy
dblp:08/4221
· DBLP profile ↗
27ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-7747-9043ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Introduction to the Special Section on software engineering for hybrid quantum computing systems
Paolo Arcaini, Andriy V. Miranskyy, Hausi A. Müller |
J. Syst. Softw. | 2 |
| 2025 | A Reference Architecture for Governance of Cloud Native ApplicationsabstractThe evolution of cloud computing has given rise to Cloud Native Applications (CNAs), presenting new challenges in governance, particularly when faced with strict compliance requirements. This work explores the unique characteristics of CNAs and their impact on governance. We introduce a comprehensive reference architecture designed to streamline governance across CNAs, along with a sample implementation, offering insights for both single and multi-cloud environments. Our architecture seamlessly integrates governance within the CNA framework, adhering to a “battery-included” philosophy. Tailored for both expansive and compact CNA deployments across various industries, this design enables cloud practitioners to prioritize product development by alleviating the complexities associated with governance. In addition, it provides a building block for academic exploration of generic CNA frameworks, highlighting their relevance in the evolving cloud computing landscape. William Pourmajidi, Lei Zhang 0078, John Steinbacher, Tony Erwin, Andriy V. Miranskyy |
IEEE Trans. Cloud Comput. | 5 |
| 2025 | Quantum Software Engineering: Roadmap and Challenges AheadabstractAs quantum computers advance, the complexity of the software they can execute increases as well. To ensure this software is efficient, maintainable, reusable, and cost-effective—key qualities of any industry-grade software—mature software engineering practices must be applied throughout its design, development, and operation. However, the significant differences between classical and quantum software make it challenging to directly apply classical software engineering methods to quantum systems. This challenge has led to the emergence of Quantum Software Engineering (QSE) as a distinct field within the broader software engineering landscape. In this work, a group of active researchers analyze in depth the current state of QSE research. From this analysis, the key areas of QSE are identified and explored in order to determine the most relevant open challenges that should be addressed in the next years. These challenges help identify necessary breakthroughs and future research directions for advancing QSE. Juan Manuel Murillo, José García-Alonso, Enrique Moguel, Johanna Barzen, Frank Leymann, Shaukat Ali 0001, Tao Yue 0002, Paolo Arcaini, Ricardo Pérez-Castillo, Ignacio García Rodríguez de Guzmán, Mario Piattini, Antonio Ruiz Cortés, Antonio Brogi, Jianjun Zhao 0001, Andriy V. Miranskyy, Manuel Wimmer |
ACM Trans. Softw. Eng. Methodol. | 15 |
| 2025 | Contrasting the Hyperparameter Tuning Impact Across Software Defect Prediction ScenariosabstractSoftware defect prediction (SDP) is crucial for delivering high-quality software products. The SDP activities help software teams better utilize their software quality assurance efforts, improving the quality of the final product. Recent research has indicated that prediction performance improvements in SDP are achievable by applying hyperparameter tuning to a particular SDP scenario (e.g., predicting defects for a future version). However, the positive impact resulting from the hyperparameter tuning step may differ based on the targeted SDP scenario. Comparing the impact of hyperparameter tuning across two SDP scenarios is necessary to provide comprehensive insights and enhance the robustness, generalizability, and, eventually, the practicality of SDP modeling for quality assurance.Therefore, in this study, we contrast the impact of hyperparameter tuning across two pivotal and consecutive SDP scenarios: (1) Inner Version Defect Prediction (IVDP) and (2) Cross Version Defect Prediction (CVDP). The main distinctions between the two scenarios lie in the scope of defect prediction and the selected evaluation setups. This study’s experiments use common evaluation setups, 28 machine learning (ML) algorithms, 53 post-release software datasets, two tuning algorithms, and five optimization metrics. We apply statistical analytics to compare the SDP performance impact differences by investigating the overall impact, the single ML algorithm impact, and variations across different software dataset sizes.The results indicate that the SDP gains within the IVDP scenario are significantly larger than those within the CVDP scenario. The results reveal that asserting performance gains for up to 24 out of 28 ML algorithms may not hold across multiple SDP scenarios. Furthermore, we found that small software datasets are more susceptible to larger differences in performance impacts. Overall, the study findings recommend software engineering researchers and practitioners to consider the effect of the selected SDP scenario when expecting performance gains from hyperparameter tuning. Mohamed Sami Rakha, Andriy V. Miranskyy, Daniel Alencar da Costa |
IEEE Trans. Software Eng. | 2 |
| 2023 | Identifying Flakiness in Quantum ProgramsabstractIn recent years, software engineers have explored ways to assist quantum software programmers. Our goal in this paper is to continue this exploration and see if quantum software programmers deal with some problems plaguing classical programs. Specifically, we examine whether intermittently failing tests, i.e., flaky tests, affect quantum software development. To explore flakiness, we conduct a preliminary analysis of 14 quantum software repositories. Then, we identify flaky tests and categorize their causes and methods of fixing them. We find flaky tests in 12 out of 14 quantum software repositories. In these 12 repositories, the lower boundary of the percentage of issues related to flaky tests ranges between 0.26% and 1.85% per repository. We identify 46 distinct flaky test reports with 8 groups of causes and 7 common solutions. Further, we notice that quantum programmers are not using some of the recent flaky test countermeasures developed by software engineers. This work may interest practitioners, as it provides useful insight into the resolution of flaky tests in quantum programs. Researchers may also find the paper helpful as it offers quantitative data on flaky tests in quantum software and points to new research opportunities. Lei Zhang 0078, Mahsa Radnejad, Andriy V. Miranskyy |
ESEM | 3 |
| 2023 | Making existing software quantum safe: A case study on IBM Db2
Lei Zhang 0078, Andriy V. Miranskyy, Walid Rjaibi, Greg Stager, Michael Gray, John Peck |
Inf. Softw. Technol. | 2 |
| 2023 | Automated data validation: An industrial experience report
Lei Zhang 0078, Sean Howard, Tom Montpool, Jessica Moore, Krittika Mahajan, Andriy V. Miranskyy |
J. Syst. Softw. | 6 |
| 2023 | Immutable Log Storage as a Service on Private and Public BlockchainsabstractService Level Agreements (SLA) are employed to ensure the performance of Cloud solutions. When a component fails, the importance of logs increases significantly. All departments may turn to logs to determine the cause of the issue and find the party at fault. The party at fault may be motivated to tamper with the logs to hide their role. We argue that the critical nature of Cloud logs calls for immutability and verification mechanism without the presence of a single trusted party. This article proposes such a mechanism by describing a blockchain-based log storage system, called Logchain, which can be integrated with existing private and public blockchain solutions. Logchain uses the immutability feature of blockchain to provide a tamper-resistance platform for log storage. Additionally, we propose a hierarchical structure to address blockchains’ scalability issues. To validate the mechanism, we integrate Logchain into Ethereum and IBM Blockchain. We show that the solution is scalable and perform the analysis of the cost of ownership to help a reader select an implementation that would address their needs. The Logchain's scalability improvement on a blockchain is achieved without any alteration of blockchains’ fundamental architecture. As shown in this work, it can function on private and public blockchains and, therefore, can be a suitable alternative for organizations that need a secure, immutable log storage platform. William Pourmajidi, Lei Zhang 0078, John Steinbacher, Tony Erwin, Andriy V. Miranskyy |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Term interrelations and trends in software engineeringabstractThe Software Engineering (SE) community is prolific, making it challenging for experts to keep up with the flood of new papers and for neophytes to enter the field. Therefore, we posit that the community may benefit from a tool extracting terms and their interrelations from the SE community's text corpus and showing terms' trends. In this paper, we build a prototyping tool using the word embedding technique. We train the embeddings on the SE Body of Knowledge handbook and 15,233 research papers' titles and abstracts. We also create test cases necessary for validation of the training of the embeddings. We provide representative examples showing that the embeddings may aid in summarizing terms and uncovering trends in the knowledge base. Janusan Baskararajah, Lei Zhang 0078, Andriy V. Miranskyy |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Anomaly Detection in Cloud ComponentsabstractCloud platforms, under the hood, consist of a complex inter-connected stack of hardware and software components. Each of these components can fail which may lead to an outage. Our goal is to improve the quality of Cloud services through early detection of such failures by analyzing resource utilization metrics. We tested Gated-Recurrent-Unit-based autoencoder with a likelihood function to detect anomalies in various multidimensional time series and achieved high performance. Mohammad Saiful Islam, Andriy V. Miranskyy |
CLOUD | 2 |
| 2018 | Logchain: Blockchain-Assisted Log StorageabstractDuring the normal operation of a Cloud solution, no one usually pays attention to the logs except technical department, which may periodically check them to ensure that the performance of the platform conforms to the Service Level Agreements. However, the moment the status of a component changes from acceptable to unacceptable, or a customer complains about accessibility or performance of a platform, the importance of logs increases significantly. Depending on the scope of the issue, all departments, including management, customer support, and even the actual customer, may turn to logs to find out what has happened, how it has happened, and who is responsible for the issue. The party at fault may be motivated to tamper the logs to hide their fault. Given the number of logs that are generated by the Cloud solutions, there are many tampering possibilities. While tamper detection solution can be used to detect any changes in the logs, we argue that critical nature of logs calls for immutability. In this work, we propose a blockchain-based log system, called Logchain, that collects the logs from different providers and avoids log tampering by sealing the logs cryptographically and adding them to a hierarchical ledger, hence, providing an immutable platform for log storage. William Pourmajidi, Andriy V. Miranskyy |
IEEE CLOUD | 2 |
| 2018 | Architecture for Analysis of Streaming DataabstractWhile several attempts have been made to construct a scalable and flexible architecture for analysis of streaming data, no general model to tackle this task exists. Thus, our goal is to build a scalable and maintainable architecture for performing analytics on streaming data. To reach this goal, we introduce a 7-layered architecture consisting of microservices and publish-subscribe software. Our study shows that this architecture yields a good balance between scalability and maintainability due to high cohesion and low coupling of the solution, as well as asynchronous communication between the layers. This architecture can help practitioners to improve their analytic solutions. It is also of interest to academics, as it is a building block for a general architecture for processing streaming data. Sheik Hoque, Andriy V. Miranskyy |
IC2E | 2 |
| 2018 | On the use of hidden Markov model to predict the time to fix bugsabstractA significant amount of time is spent by software developers in investigating bug reports. It is useful to indicate when a bug report will be closed, since it would help software teams to prioritise their work. Several studies have been conducted to address this problem in the past decade. Most of these studies have used the frequency of occurrence of certain developer activities as input attributes in building their prediction models. However, these approaches tend to ignore the temporal nature of the occurrence of these activities. In this paper, a novel approach using Hidden Markov models (HMMs) and temporal sequences of developer activities is proposed. The approach is empirically demonstrated in a case study using eight years of bug reports collected from the Firefox project. We provide additional details below. In a software bug repository, recorded developer activities occur sequentially. For example, activity C (a certain person has been copied on the bug report) is followed by activity A (bug confirmed and assigned to a named developer), which in turn is followed by activity Z (bug reached status resolved). Additional piece of information is developers' level of expertise, such as novice (N), intermediate (M), or experienced (E), at the time of report creation. We combine these data together to produce a sequence of temporal activities associated with bug reports in the Firefox bug repository. Mayy Habayeb, Syed Shariyar Murtaza, Andriy V. Miranskyy, Ayse Basar Bener |
ICSE | 3 |
| 2018 | Editorial: Special Section on Best Papers of PROMISE 2016
Hongyu Zhang 0002, Andriy V. Miranskyy, Ayse Basar Bener |
Inf. Softw. Technol. | 2 |
| 2018 | Database engines: Evolution of greennessabstractAbstract Information technology consumes up to 10% of the world's electricity generation, contributing to CO2emissions and high energy costs. Data centers, particularly databases, use up to 23% of this energy. Therefore, building an energy‐efficient (green) database engine could reduce energy consumption and CO2emissions. The goal of this study is to understand the factors driving databases' energy consumption and execution time throughout their evolution. We conducted an empirical case study of energy consumption by 2 MySQL database engines, InnoDB and MyISAM, across 40 releases. We examined the relationships of 4 software metrics to energy consumption and execution time to determine which metrics reflect the greenness and performance of a database. Our analysis shows that database engines' energy consumption and execution time increase as databases evolve. Moreover, the lines of code (LOC) metric is correlated moderately to strongly with energy consumption and execution time in 88% of cases. Our findings provide insights to practitioners and researchers. Database administrators may use them to select a fast, green release of the MySQL database engine. MySQL developers may use LOC to assess products' greenness and performance. Researchers may use our findings to further develop new hypotheses or build models predicting greenness and performance of databases. Andriy V. Miranskyy, Zainab Al-Zanbouri, David Godwin, Ayse Basar Bener |
J. Softw. Evol. Process. | 1 |
| 2018 | On the Use of Hidden Markov Model to Predict the Time to Fix BugsabstractA significant amount of time is spent by software developers in investigating bug reports. It is useful to indicate when a bug report will be closed, since it would help software teams to prioritise their work. Several studies have been conducted to address this problem in the past decade. Most of these studies have used the frequency of occurrence of certain developer activities as input attributes in building their prediction models. However, these approaches tend to ignore the temporal nature of the occurrence of these activities. In this paper, a novel approach using Hidden Markov Models and temporal sequences of developer activities is proposed. The approach is empirically demonstrated in a case study using eight years of bug reports collected from the Firefox project. Our proposed model correctly identifies bug reports with expected bug fix times. We also compared our proposed approach with the state of the art technique in the literature in the context of our case study. Our approach results in approximately 33 percent higher F-measure than the contemporary technique based on the Firefox project data. Mayy Habayeb, Syed Shariyar Murtaza, Andriy V. Miranskyy, Ayse Basar Bener |
IEEE Trans. Software Eng. | 3 |
| 2017 | Rediscovery datasets: connecting duplicate reportsabstractThe same defect can be rediscovered by multiple clients, causing unplanned outages and leading to reduced customer satisfaction. In the case of popular open source software, high volume of defects is reported on a regular basis. A large number of these reports are actually duplicates / rediscoveries of each other. Researchers have analyzed the factors related to the content of duplicate defect reports in the past. However, some of the other potentially important factors, such as the inter-relationships among duplicate defect reports, are not readily available in defect tracking systems such as Bugzilla. This information may speed up bug fixing, enable efficient triaging, improve customer profiles, etc. In this paper, we present three defect rediscovery datasets mined from Bugzilla. The datasets capture data for three groups of open source software projects: Apache, Eclipse, and KDE. The datasets contain information about approximately 914 thousands of defect reports over a period of 18 years (1999-2017) to capture the inter-relationships among duplicate defects. We believe that sharing these data with the community will help researchers and practitioners to better understand the nature of defect rediscovery and enhance the analysis of defect reports. Mefta Sadat, Ayse Basar Bener, Andriy V. Miranskyy |
MSR | 3 |
| 2015 | Merits of Organizational Metrics in Defect Prediction: An Industrial ReplicationabstractDefect prediction models presented in the literature lack generalization unless the original study can be replicated using new datasets and in different organizational settings. Practitioners can also benefit from replicating studies in their own environment by gaining insights and comparing their findings with those reported. In this work, we replicated an earlier study in order to investigate the merits of organizational metrics in building defect prediction models for large-scale enterprise software. We mined the organizational, code complexity, code churn and pre-release bug metrics of that large scale software and built defect prediction models for each metric set. In the original study, organizational metrics were found to achieve the highest performance. In our case, models based on organizational metrics performed better than models based on churn metrics but were outperformed by pre-release metric models. Further, we verified four individual organizational metrics as indicators for defects. We conclude that the performance of different metric sets in building defect prediction models depends on the project's characteristics and the targeted prediction level. Our replication of earlier research enabled assessing the validity and limitations of organizational metrics in a different context. Bora Caglayan, Burak Turhan, Ayse Basar Bener, Mayy Habayeb, Andriy V. Miranskyy, Enzo Cialini |
ICSE (2) | 5 |
| 2015 | 4th International Workshop on Realizing AI Synergies in Software Engineering (RAISE 2015)abstractThis workshop is the fourth in the series and continued to build upon the work carried out at the previous iterations of the International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering, which were held at ICSE in 2012, 2013 and 2014. RAISE 2015 brought together researchers and practitioners from the artificial intelligence (AI) and software engineering (SE) disciplines to build on the interdis- ciplinary synergies that exist and to stimulate further interaction across these disciplines. Mutually beneficial characteristics have appeared in the past few decades and are still evolving due to new challenges and technological advances. Hence, the question that motivates and drives the RAISE Workshop series is: "Are SE and AI researchers ignoring important insights from AI and SE?". To pursue this question, RAISE'15 explored not only the application of AI techniques to SE problems but also the application of SE techniques to AI problems. RAISE not only strengthens the AI- and-SE community but also continues to develop a roadmap of strategic research directions for AI and SE. Burak Turhan, Ayse Basar Bener, Rachel Harrison, Andriy V. Miranskyy, Çetin Meriçli, Leandro L. Minku |
ICSE (2) | 4 |
| 2015 | The Firefox Temporal Defect DatasetabstractThe bug tracking repositories of software projects capture initial defect (bug) reports and the history of interactions among developers, testers, and customers. Extracting and mining information from these repositories is time consuming and daunting. Researchers have focused mostly on analyzing the frequency of the occurrence of defects and their attributes (e.g., The number of comments and lines of code changed, count of developers). However, the counting process eliminates information about the temporal alignment of events leading to changes in the attributes count. Software quality teams could plan and prioritize their work more efficiently if they were aware of these temporal sequences and knew their frequency of occurrence. In this paper, we introduce a novel dataset mined from the Fire fox bug repository (Bugzilla) which contains information about the temporal alignment of developer interactions. Our dataset covers eight years of data from the Fire fox project on activities throughout the project's lifecycle. Some of these activities have not been reported in frequency-based or other temporal datasets. The dataset we mined from the Fire fox project contains new activities, such as reporter experience, file exchange events, code-review process activities, and setting of milestones. We believe that this new dataset will improve analysis of bug reports and enable mining of temporal relationships so that practitioners can enhance their bug-fixing process. Mayy Habayeb, Andriy V. Miranskyy, Syed Shariyar Murtaza, Leotis Buchanan, Ayse Basar Bener |
MSR | 2 |
| 2015 | Predicting defective modules in different test phases
Bora Caglayan, Ayse Tosun Misirli, Ayse Basar Bener, Andriy V. Miranskyy |
Softw. Qual. J. | 4 |
| 2014 | Effect of temporal collaboration network, maintenance activity, and experience on defect exposureabstractContext: Number of defects fixed in a given month is used as an input for several project management decisions such as release time, maintenance effort estimation and software quality assessment. Past activity of developers and testers may help us understand the future number of reported defects. Goal: To find a simple and easy to implement solution, predicting defect exposure. Method: We propose a temporal collaboration network model that uses the history of collaboration among developers, testers, and other issue originators to estimate the defect exposure for the next month. Results: Our empirical results show that temporal collaboration model could be used to predict the number of exposed defects in the next month with R2 values of 0.73. We also show that temporality gives a more realistic picture of collaboration network compared to a static one. Conclusions: We believe that our novel approach may be used to better plan for the upcoming releases, helping managers to make evidence based decisions. Andriy V. Miranskyy, Bora Caglayan, Ayse Basar Bener, Enzo Cialini |
ESEM | 1 |
| 2012 | Using entropy measures for comparison of software traces
Andriy V. Miranskyy, Matthew Davison 0001, Mark Reesor, Syed Shariyar Murtaza |
Inf. Sci. | 1 |
| 2011 | Characteristics of multiple-component defects and architectural hotspots: a large system case study
Zude Li, Nazim H. Madhavji, Syed Shariyar Murtaza, Mechelle Gittens, Andriy V. Miranskyy, David Godwin, Enzo Cialini |
Empir. Softw. Eng. | 5 |
| 2009 | Analysis of pervasive multiple-component defects in a large software systemabstractCertain software defects require corrective changes repeatedly in a few components of the system. One type of such defects spans multiple components of the system, and we call such defects pervasive multiple-component defects (PMCDs). In this paper, we describe an empirical study of six releases of a large legacy software system (of approx. size 20 million physical lines of code) to analyze PMCDs with respect to: (1) the complexity of fixing such defects and (2) the persistence of defect-prone components across phases and releases. The overall hypothesis in this study is that PMCDs inflict a greater negative impact than do other defects on defect-correction efficacy. Our findings show that the average number of changes required for fixing PMCDs is 20-30 times as much as the average for all defects. Also, over 80% of PMCD-contained defect-prone components still remain defect-prone in successive phases or releases. These findings support the overall hypothesis strongly. We compare our results, where possible, to those of other researchers and discuss the implications on maintenance processes and tools. Zude Li, Mechelle Gittens, Syed Shariyar Murtaza, Nazim H. Madhavji, Andriy V. Miranskyy, David Godwin, Enzo Cialini |
ICSM | 5 |
| 2007 | An iterative, multi-level, and scalable approach to comparing execution tracesabstractIn this paper, we overview a new approach to comparing execution traces. Such comparison can be useful for purposes such as improving test coverage and profiling system's users. In our approach, traces are compressed into different levels of compaction and are then compared iteratively from highest to lowest levels, rejecting dissimilar traces in the process and eventually leaving residual, similar traces. These residual traces form an important feedback for improvement or analysis goals. The preliminary results show that the approach is scalable for industrial use. Andriy V. Miranskyy, Nazim H. Madhavji, Mechelle Gittens, Matthew Davison 0001, Mark Wilding, David Godwin |
ESEC/SIGSOFT FSE | 1 |
| 2005 | Modelling Assumptions and Requirements in the Context of Project RiskabstractMany researchers have emphasized the importance of documenting assumptions (As) underlying software requirements (Rs). However, As and Rs can change with time for reasons such as: (i) an A or R was elicited incorrectly and subsequently needs to be changed; (ii) operational domain changes induce changes in the A and R sets; and (iii) the change in validity of an A, or desirability of an R, respectively, causes the validity of another A or desirability of an R to change. In Section 2, we describe our model and how it works. To put such a model into practice, we need to consider at least two scenarios. One is intra-release cycle-time, where invalidity risk is predicted at the start of the project for times between the inception and completion of the project. This would give us intra-release risk trends. The second scenario is prediction over multiple releases. This would give us a risk trend over a longer period of time. The full paper describes an algorithm to cover both of these scenarios and gives an example (from a banking application) of how the model could apply in practice. Here, we consider only the first scenario due to limitation of space. Andriy V. Miranskyy, Nazim H. Madhavji, Matthew Davison 0001, Mark Reesor |
RE | 1 |