Cor-Paul Bezemer

dblp:95/7251 · DBLP profile ↗
← Back
74ranked-venue papers
8as first author
36since 2021 · last 2027
0000-0002-0474-5718ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 63 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2027 Exploring the capabilities of vision-language models to detect visual bugs in HTML5 canvas applications
Finlay Macklon, Cedric Boucher, Cor-Paul Bezemer
Empir. Softw. Eng.3
2026 A systematic literature review of software engineering research on Jupyter notebook
abstract
• This research provides the first comprehensive systematic literature review on software engineering research specifically targeting Jupyter notebooks, identifying 199 primary studies published up to September 2025 and categorizing them into 11 core software engineering topics. • This research reveals that a large portion of the studies have been published outside traditional software engineering venues, with Human-Computer Interaction conferences like ACM Conference on Human Factors in Computing Systems (CHI) being the top publishing venues, highlighting the interdisciplinary nature of Jupyter Notebook research. • This research identifies a reusability gap in existing research, showing that only 82 out of 199 studies offer usable replication packages, and most are hosted on GitHub instead of permanent repositories, which violates open science best practices. • This research identifies that notebook-specific solutions for software engineering issues such as testing, refactoring, and documentation are relatively underexplored. Future directions include resolving duplicated execution numbers, refactoring inter-notebook clones, and generating grouped documentation for coherent-code cells are future directions derived from our study. • This research proposes the integration of modern AI-based solutions into Jupyter notebooks to support various software engineering topics, including code search and code generation. Additionally, future research should leverage advanced AI techniques (e.g., large language models), to improve conversational AI-powered assistants for automated code generation by multi-step workflow automation in data science notebooks. • Although the paper exceeds the recommended length due to the inclusion of detailed tables, figures, and categorized analyses (covering 11 topics and 21 subtopics), we believe that this extended content is essential for clearly and completely reporting our findings. As the first systematic literature review in this domain, we have carefully structured the paper to ensure readability. We believe the length is justified by the value and breadth of this paper’s contributions. Context : Jupyter Notebook has emerged as a versatile tool that transforms how researchers, developers, and data scientists conduct and communicate their work. As the adoption of Jupyter notebooks continues to rise, so does the interest from the software engineering research community in improving the software engineering practices for Jupyter notebooks. Objective : The purpose of this study is to analyze trends, gaps, and methodologies used in software engineering research on Jupyter notebooks. Method : We selected 199 relevant publications up to September 2025, following established systematic literature review guidelines. We explored publication trends, categorized them based on software engineering topics, and reported findings based on those topics. Results : The most popular venues for publishing software engineering research on Jupyter notebooks are related to human-computer interaction instead of traditional software engineering venues. Researchers have addressed a wide range of software engineering topics on notebooks, such as code reuse, readability, and execution environment. Although reusability is one of the research topics for Jupyter notebooks, only 82 of the 199 studies can be reused based on their provided URLs. Additionally, most replication packages are not hosted on permanent repositories for long-term availability and adherence to open science principles. Conclusion : Solutions specific to notebooks for software engineering issues, including testing, refactoring, and documentation, are underexplored. Future research opportunities exist in automatic testing frameworks, refactoring clones between notebooks, and generating group documentation for coherent code cells.
Md. Saeed Siddik, Hao Li 0094, Cor-Paul Bezemer
J. Syst. Softw.3
2025 Early Detection of Performance Regressions by Bridging Local Performance Data and Architectural Models
abstract
During software development, developers often make numerous modifications to the software to address existing issues or implement new features. However, certain changes may inadvertently have a detrimental impact on the overall system performance. To ensure that the performance of new software re-leases does not degrade (i.e., absence of performance regressions), existing practices rely on system-level performance testing, such as load testing, or component-level performance testing, such as microbenchmarking, to detect performance regressions. However, performance testing for the entire system is often expensive and time-consuming, posing challenges to adapting to the rapid release cycles common in modern DevOps practices. In addition, system-level performance testing cannot be conducted until the system is fully built and deployed. On the other hand, component-level testing focuses on isolated components, neglecting overall system performance and the impact of system workloads. In this paper, we propose a novel approach to early detection of performance regressions by bridging the local performance data generated by component-level testing and the system-level architectural models. Our approach uses local performance data to identify deviations at the component level, and then propagate these deviations to the architectural model. We then use the architectural model to predict regressions in the performance of the overall system. In an evaluation of our approach on two representative open-source benchmark systems, we show that it can effectively detect end-to-end system performance regressions from local performance deviations with different intensities and under various system workloads. More importantly, our approach can detect regressions as early as in the development phase, in contrast to existing approaches that require the system to be fully built and deployed. Our approach is lightweight and can complement traditional system performance testing when testing resources are scarce.
Lizhi Liao, Simon Eismann, Heng Li 0007, Cor-Paul Bezemer, Diego Costa 0001, André van Hoorn, Weiyi Shang
ICSE4
2025 Answering User Questions About Machine Learning Models Through Standardized Model Cards
abstract
Reusing pre-trained machine learning models is becoming very popular due to model hubs such as Hugging Face (HF). However, similar to when reusing software, many issues may arise when reusing an ML model. In many cases, users resort to asking questions on discussion forums such as the HF community forum. In this paper, we study how we can reduce the community's workload in answering these questions and increase the likelihood that questions receive a quick answer. We analyze 11,278 discussions from the HF model community that contain user questions about ML models. We focus on the effort spent handling questions, the high-level topics of discussions, and the potential for standardizing responses in model cards based on a model card template. Our findings indicate that there is not much effort involved in responding to user questions, however, 40.1% of the questions remain open without any response. A topic analysis shows that discussions are more centered around technical details on model development and troubleshooting, indicating that more input from model providers is required. We show that 42.5% of the questions could have been answered if the model provider followed a standard model card template for the model card. Based on our analysis, we recommend that model providers add more development-related details on the model's architecture, algorithm, data preprocessing and training code in existing documentation (sub)sections and add new (sub)sections to the template to address common questions about model usage and hardware requirements.
Tajkia Rahman Toma, Balreet Grewal, Cor-Paul Bezemer
ICSE3
2025 VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
abstract
With video games leading in entertainment revenues, optimizing game development workflows is critical to the industry’s long-term success. Recent advances in vision-language models (VLMs) hold significant potential to automate and enhance various aspects of game development—particularly video game quality assurance (QA), which remains one of the most labor-intensive processes with limited automation. To effectively measure VLM performance in video game QA tasks and evaluate their ability to handle real-world scenarios, there is a clear need for standardized benchmarks, as current ones fall short in addressing this domain. To bridge this gap, we introduce VideoGameQA-Bench - a comprehensive benchmark designed to encompass a wide range of game QA activities, including visual unit testing, visual regression testing, needle-in-a-haystack, glitch detection, and bug report generation for both images and videos.
Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zad Tootaghaj, Nabajeet Barman, Cor-Paul Bezemer
NeurIPS5
2025 Videogamebunny: Towards Vision Assistants for Video Games
abstract
Large multimodal models (LMMs) hold substantial promise across various domains, from personal assistance in daily tasks to sophisticated applications like medical diagnostics. However, their capabilities have limitations in the video game domain, such as challenges with scene understanding, hallucinations, and inaccurate descriptions of video game content, especially in open-source models. This paper describes the development of VideoGameBunny,a LLaVA-style model based on Bunny, specifically tailored for understanding images from video games. We release intermediate checkpoints, training logs, and an extensive dataset comprising 185,259 video game images from 413 titles, along with 389,565 image-instruction pairs that include image captions, question-answer pairs, and a JSON representation of 16 elements of 136,974 images. Our experiments show that our high quality game-related data has the potential to make a relatively small model outperform the much larger state-of-the-art model LLaVa-1.6-34b (which has more than 4x the number of parameters). Our study paves the way for future research in video game understanding on tasks such as playing, commentary, and debugging. Code and data are available at: https://videogamebunny.github.io/
Mohammad Reza Taesiri, Cor-Paul Bezemer
WACV2
2025 Bridging the language gap: an empirical study of bindings for open source machine learning libraries across software package ecosystems
Hao Li 0094, Cor-Paul Bezemer
Empir. Softw. Eng.2
2025 Detecting Discrepancies Between Subtitles and Audio in Gameplay Videos With EchoTest
abstract
The landscape of accessibility features in video games remains inconsistent, posing challenges for gamers who seek experiences tailored to their needs. Accessibility features, such as subtitles are widely used by players but are difficult to test manually due to the large scope of games and the variability in how subtitles can appear. In this article, we introduce an automated approach (EchoTest) to extract subtitles and spoken audio from a gameplay video, convert them into text, and compare them to detect discrepancies, such as typos, desynchronization, and missing text.EchoTestcan be used by game developers to identify discrepancies between subtitles and spoken audio in their games, enabling them to better test the accessibility of their games. In an empirical study on gameplay videos from 15 popular games,EchoTestcan verify discrepancies between subtitles and audio with a precision of 98% and a recall of 89%. In addition,EchoTestperforms well with a precision of 73% and a recall of 99% on a challenging generated benchmark.
Ian Gauk, Cor-Paul Bezemer
IEEE Trans. Games2
2025 Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
abstract
Bindings for machine learning frameworks (such as TensorFlow and PyTorch) allow developers to integrate a framework’s functionality using a programming language different from the framework’s default language (usually Python). In this article, we study the impact of using TensorFlow and PyTorch bindings in C#, Rust, Python and JavaScript on the software quality in terms of correctness (training and test accuracy) and time cost (training and inference time) when training and performing inference on five widely used deep learning models. Our experiments show that a model can be trained in one binding and used for inference in another binding for the same framework without losing accuracy. Our study is the first to show that using a non-default binding can help improve machine learning software quality from the time cost perspective compared to the default Python binding while still achieving the same level of correctness.
Hao Li 0094, Gopi Krishnan Rajbahadur, Cor-Paul Bezemer
ACM Trans. Softw. Eng. Methodol.3
2024 An Exploratory Study of Dataset and Model Management in Open Source Machine Learning Applications
abstract
Datasets and models are two key artifacts in machine learning (ML) applications. Although there exist tools to support dataset and model developers in managing ML artifacts, little is known about how these datasets and models are integrated into ML applications. In this paper, we study how datasets and models in ML applications are managed. In particular, we focus on how these artifacts are stored and versioned alongside the applications. After analyzing 93 repositories, we identified the most common storage location to store datasets and models is the file system, which causes availability issues. Notably, large data and model files, exceeding approximately 60 MB, are stored exclusively in remote storage and downloaded as needed. Most of the datasets and models lack proper integration with the version control system, posing potential trace-ability and reproducibility issues. Additionally, although datasets and models are likely to evolve during the application development, they are rarely updated in application repositories.
Tajkia Rahman Toma, Cor-Paul Bezemer
CAIN2
2024 GlitchBench: Can Large Multimodal Models Detect Video Game Glitches?
abstract
Large multimodal models (LMMs) have evolved from large language models (LLMs) to integrate multiple input modalities, such as visual inputs. This integration augments the capacity of LLMs for tasks requiring visual comprehen- sion and reasoning. However, the extent and limitations of their enhanced abilities are not fully understood, especially when it comes to real-world tasks. To address this gap, we introduce GlitchBench, a novel benchmark derived from video-game quality assurance tasks, to test and eval- uate the reasoning capabilities of LMMs. Our benchmark is curated from a variety of unusual and glitched scenarios from video games and aims to challenge both the visual and linguistic reasoning powers of LMMs in detecting and inter- preting out-of-the-ordinary events. Our evaluation shows that GlitchBench presents a new, interesting challenge to state-of-the-art LMMs. Code and data are available at: https://glitchbench.github.io/
Mohammad Reza Taesiri, Tianjun Feng, Cor-Paul Bezemer, Anh Totti Nguyen
CVPR3
2024 Analyzing Developer Use of ChatGPT Generated Code in Open Source GitHub Projects
abstract
The rapid development of large language models such as ChatGPT have made them particularly useful to developers in generating code snippets for their projects. To understand how ChatGPT's generated code is leveraged by developers, we conducted an empirical study of 3,044 ChatGPT-generated code snippets integrated within GitHub projects. A median of 54% of the generated lines of code is found in the project's code and this code typically remains unchanged once added. The modifications of the 76 code snippets that changed in a subsequent commit, consisted of minor functionality changes and code reorganizations that were made within a day. Our findings offer insights that help drive the development of AI-assisted programming tools. We highlight the importance of making changes in ChatGPT code before integrating it into a project.
Balreet Grewal, Sarah Nadi, Cor-Paul Bezemer
MSR4
2024 Searching Bug Instances in Gameplay Video Repositories
abstract
Gameplay videos offer valuable insights into player interactions and game responses, particularly data about game bugs. Despite the abundance of gameplay videos online, extracting useful information remains a challenge. This paper introduces a method for searching and extracting relevant videos from extensive video repositories using English text queries. Our approach requires no external information, like video metadata; it solely depends on video content. Leveraging the zero-shot transfer capabilities of the Contrastive Language-Image Pre-Training (CLIP) model, our approach does not require any data labeling or training. To evaluate our approach, we present theGamePhysicsdataset, comprising 26,954 videos from 1,873 games that were collected from the GamePhysics section on the Reddit website. Our approach shows promising results in our extensive analysis of simple and compound queries, indicating that our method is useful for detecting objects and events in gameplay videos. Moreover, we assess the effectiveness of our method by analyzing a carefully annotated dataset of 220 gameplay videos. The results of our study demonstrate the potential of our approach for applications such as the creation of a video search tool tailored to identifying video game bugs, which could greatly benefit Quality Assurance (QA) teams in finding and reproducing bugs. The code and data used in this paper can be found athttps://zenodo.org/records/10211390
Mohammad Reza Taesiri, Finlay Macklon, Sarra Habchi, Cor-Paul Bezemer
IEEE Trans. Games4
2024 Leveraging the OPT Large Language Model for Sentiment Analysis of Game Reviews
abstract
Automatically extracting players' sentiments about games can help game developers to better understand the aspects of their games that players like or dislike. Our prior work showed that traditional sentiment analysis techniques do not perform well on game reviews. However, the Natural Language Processing (NLP) field has seen a steep progress in recent years. In this letter, we follow up on our prior work and investigate how a state-of-the-art large language model (OPT-175B) performs on the sentiment classification of game reviews. We manually analyze the game reviews wrongly classified by OPT-175B to better understand the issues that affect the performance of that model and how those issues compare to the challenges faced by traditional classifiers. We found that OPT-175B achieves (far) better performance than traditional sentiment classifiers, with a 72%-increased F-measure and a 30%-increased AUC compared to the best traditional classifier studied in our prior work. We also found that common challenges of traditional classifiers, such as reviews with game comparisons and negative terminology, have been mostly solved by the OPT-175B model.
Markos Viggiato, Cor-Paul Bezemer
IEEE Trans. Games2
2023 ImageNet-Hard: The Hardest Images Remaining from a Study of the Power of Zoom and Spatial Biases in Image Classification
abstract
Image classifiers are information-discarding machines, by design. Yet, how these models discard information remains mysterious. We hypothesize that one way for image classifiers to reach high accuracy is to first zoom to the most discriminative region in the image and then extract features from there to predict image labels, discarding the rest of the image. Studying six popular networks ranging from AlexNet to CLIP, we find that proper framing of the input image can lead to the correct classification of 98.91% of ImageNet images. Furthermore, we uncover positional biases in various datasets, especially a strong center bias in two popular datasets: ImageNet-A and ObjectNet. Finally, leveraging our insights into the potential of zooming, we propose a test-time augmentation (TTA) technique that improves classification accuracy by forcing models to explicitly perform zoom-in operations before making predictions.Our method is more interpretable, accurate, and faster than MEMO, a state-of-the-art (SOTA) TTA method. We introduce ImageNet-Hard, a new benchmark that challenges SOTA classifiers including large vision-language models even when optimal zooming is allowed.
Mohammad Reza Taesiri, Giang Nguyen 0004, Sarra Habchi, Cor-Paul Bezemer, Anh Totti Nguyen
NeurIPS4
2023 Do Code Quality and Style Issues Differ Across (Non-)Machine Learning Notebooks? Yes!
abstract
The popularity of computational notebooks is rapidly increasing because of their interactive code-output visualization and on-demand non-sequential code block execution. These notebook features have made notebooks especially popular with machine learning developers and data scientists. However, as prior work shows, notebooks generally contain low quality code. In this paper, we investigate whether the low quality code is inherent to the programming style in notebooks, or whether it is correlated with the use of machine learning techniques. We present a large-scale empirical analysis of 246,599 opensource notebooks to explore how machine learning code quality in Jupyter Notebooks differs from non-machine learning code, thereby focusing on code style issues. We explored code style issues across the Error, Convention, Warning, and Refactoring categories. We found that machine learning notebooks are of lower quality regarding PEP-8 code standards than non-machine learning notebooks, and their code quality distributions significantly differ with a small effect size. We identified several code style issues with large differences in occurrences between machine learning and non-machine learning notebooks. For example, package and import-related issues are more prevalent in machine learning notebooks. Our study shows that code quality and code style issues differ significantly across machine learning and non-machine learning notebooks.
Md. Saeed Siddik, Cor-Paul Bezemer
SCAM2
2023 Prioritizing Natural Language Test Cases Based on Highly-Used Game Features
abstract
Software testing is still a manual activity in many industries, such as the gaming industry. But manually executing tests becomes impractical as the system grows and resources are restricted, mainly in a scenario with short release cycles. Test case prioritization is a commonly used technique to optimize the test execution. However, most prioritization approaches do not work for manual test cases as they require source code information or test execution history, which is often not available in a manual testing scenario. In this paper, we propose a prioritization approach for manual test cases written in natural language based on the tested application features (in particular, highly-used application features). Our approach consists of (1) identifying the tested features from natural language test cases (with zero-shot classification techniques) and (2) prioritizing test cases based on the features that they test. We leveraged the NSGA-II genetic algorithm for the multi-objective optimization of the test case ordering to maximize the coverage of highly-used features while minimizing the cumulative execution time. Our findings show that we can successfully identify the application features covered by test cases using an ensemble of pre-trained models with strong zero-shot capabilities (an F-score of 76.1%). Also, our prioritization approaches can find test case orderings that cover highly-used application features early in the test execution while keeping the time required to execute test cases short. QA engineers can use our approach to focus the test execution on test cases that cover features that are relevant to users.
Markos Viggiato, Dale Paas, Cor-Paul Bezemer
ESEC/SIGSOFT FSE3
2023 Analyzing Techniques for Duplicate Question Detection on Q&A Websites for Game Developers
Arthur V. Kamienski, Abram Hindle, Cor-Paul Bezemer
Empir. Softw. Eng.3
2023 Towards reducing the time needed for load testing
abstract
Abstract The performance of large‐scale systems must be thoroughly tested under various levels of workload, as load‐related issues can have a disastrous impact on the system. However, load testing often requires a large amount of time, running from hours to even days. In our prior work, we reduced the execution time of a load test by detecting repetitiveness in individual performance metric values, such as CPU utilization, that are observed during the test. However, as we explain in this paper, disregarding combinations of performance metrics may miss important information about the load‐related behavior of a system. In this paper we revisit our prior approach, by proposing an approach that reduces the execution time of a load test by detecting whether a test no longer exercises new combinations of the observed performance metrics. We study three open source systems, in which we use our new and prior approaches to reduce the execution time of a 24‐hour load test. We show that our new approach is capable of reducing the execution time of the test to less than 8.5 hours, while preserving a coverage of at least 95% of the combinations that are observed between the performance metrics during the 24‐hour tests.
Hammam M. AlGhamdi, Cor-Paul Bezemer, Weiyi Shang, Ahmed E. Hassan, Parminder Flora
J. Softw. Evol. Process.2
2023 An Empirical Study of Yanked Releases in the Rust Package Registry
abstract
Cargo, the software packaging manager of Rust, provides a yank mechanism to support release-level deprecation, which can prevent packages from depending on yanked releases. Most prior studies focused on code-level (i.e., deprecated APIs) and package-level deprecation (i.e., deprecated packages). However, few studies have focused on release-level deprecation. In this study, we investigate how often and how the yank mechanism is used, the rationales behind its usage, and the adoption of yanked releases in the Cargo ecosystem. Our study shows that 9.6% of the packages in Cargo have at least one yanked release, and the proportion of yanked releases kept increasing from 2014 to 2020. Package owners yank releases for other reasons than withdrawing a defective release, such as fixing a release that does not follow semantic versioning or indicating a package is removed or replaced. In addition, we found that 46% of the packages directly adopted at least one yanked release and the yanked releases propagated through the dependency network, which leads to 1.4% of the releases in the ecosystem having unresolved dependencies.
Hao Li 0094, Filipe Roseiro Côgo, Cor-Paul Bezemer
IEEE Trans. Software Eng.3
2023 A Taxonomy of Testable HTML5 Canvas Issues
abstract
The HTML5canvas> is widely used to display high quality graphics in web applications. However, the combination of web, GUI, and visual techniques that are required to buildcanvas> applications, together with the lack of testing and debugging tools, makes developing such applications very challenging. To help direct future research on testingcanvas> applications, in this paper we present a taxonomy of testablecanvas> issues. First, we extracted 2,403canvas>-related issue reports from 123 open source GitHub projects that use the HTML5canvas>. Second, we constructed our taxonomy by manually classifying a random sample of 332 issue reports. Our manual classification identified five broad categories of testablecanvas> issues, such as Visual and Performance issues. We found that Visual issues are the most frequent (35%), while Performance issues are relatively infrequent (5%). We also found that many testablecanvas> issues that present themselves visually on thecanvas> are actually caused by other components of the web application. Our taxonomy of testablecanvas> issues can be used to steer future research intocanvas> issues and testing.
Finlay Macklon, Markos Viggiato, Natalia Romanova, Chris Buzon, Dale Paas, Cor-Paul Bezemer
IEEE Trans. Software Eng.6
2023 Identifying Similar Test Cases That Are Specified in Natural Language
abstract
Software testing is still a manual process in many industries, despite the recent improvements in automated testing techniques. As a result, test cases (which consist of one or more test steps that need to be executed manually by the tester) are often specified in natural language by different employees and many redundant test cases might exist in the test suite. This increases the (already high) cost of test execution. Manually identifying similar test cases is a time-consuming and error-prone task. Therefore, in this paper, we propose an unsupervised approach to identify similar test cases. Our approach uses a combination of text embedding, text similarity and clustering techniques to identify similar test cases. We evaluate five different text embedding techniques, two text similarity metrics, and two clustering techniques to cluster similar test steps and three techniques to identify similar test cases from the test step clusters. Through an evaluation in an industrial setting, we showed that our approach achieves a high performance to cluster test steps (an F-score of 87.39%) and identify similar test cases (an F-score of 86.13%). Furthermore, a validation with developers indicates several different practical usages of our approach (such as identifying redundant test cases), which help to reduce the testing manual effort and time.
Markos Viggiato, Dale Paas, Chris Buzon, Cor-Paul Bezemer
IEEE Trans. Software Eng.4
2022 How are Solidity smart contracts tested in open source projects? An exploratory study
abstract
Smart contracts are self-executing programs that are stored on the blockchain. Once a smart contract is compiled and deployed on the blockchain, it cannot be modified. Therefore, having a bug-free smart contract is vital. To ensure a bug-free smart contract, it must be tested thoroughly. However, little is known about how developers test smart contracts in practice. Our study explores 139 open source smart contract projects that are written in Solidity to investigate the state of smart contract testing from three dimensions: (1) the developers working on the tests, (2) the used testing frameworks and testnets and (3) the type of tests that are conducted. We found that mostly core developers of a project are responsible for testing the contracts. Second, developers typically use only functional testing frameworks to test a smart contract, with Truffle being the most popular one. Finally, our results show that functional testing is conducted in most of the studied projects (93%), security testing is only performed in a few projects (9.4%) and traditional performance testing is conducted in none. In addition, we found 34 projects that mentioned or published external audit reports.
Luisa Palechor, Cor-Paul Bezemer
AST2
2022 Automatically Detecting Visual Bugs in HTML5 Canvas Games
abstract
The HTML5 is used to display high quality graphics in web applications such as web games (i.e., games). However, automatically testing games is not possible with existing web testing techniques and tools, and manual testing is laborious. Many widely used web testing tools rely on the Document Object Model (DOM) to drive web test automation, but the contents of the are not represented in the DOM. The main alternative approach, snapshot testing, involves comparing oracle snapshot images with test-time snapshot images using an image similarity metric to catch visual bugs, i.e., bugs in the graphics of the web application. However, creating and maintaining oracle snapshot images for games is onerous, defeating the purpose of test automation. In this paper, we present a novel approach to automatically detect visual bugs in games. By leveraging an internal representation of objects on the , we decompose snapshot images into a set of object images, each of which is compared with a respective oracle asset (e.g., a sprite) using four similarity metrics: percentage overlap, mean squared error, structural similarity, and embedding similarity. We evaluate our approach by injecting 24 visual bugs into a custom game, and find that our approach achieves an accuracy of 100%, compared to an accuracy of 44.6% with traditional snapshot testing.
Finlay Macklon, Mohammad Reza Taesiri, Markos Viggiato, Stefan Antoszko, Natalia Romanova, Dale Paas, Cor-Paul Bezemer
ASE7
2022 CLIP meets GamePhysics: Towards bug identification in gameplay videos using zero-shot transfer learning
abstract
Gameplay videos contain rich information about how players interact with the game and how the game responds. Sharing gameplay videos on social media platforms, such as Reddit, has become a common practice for many players. Often, players will share game-play videos that showcase video game bugs. Such gameplay videos are software artifacts that can be utilized for game testing, as they provide insight for bug analysis. Although large repositories of gameplay videos exist, parsing and mining them in an effective and structured fashion has still remained a big challenge. In this paper, we propose a search method that accepts any English text query as input to retrieve relevant videos from large repositories of gameplay videos. Our approach does not rely on any external information (such as video metadata); it works solely based on the content of the video. By leveraging the zero-shot transfer capabilities of the Contrastive Language-Image Pre-Training (CLIP) model, our approach does not require any data labeling or training. To evaluate our approach, we present the GamePhysics dataset consisting of 26,954 videos from 1,873 games, that were collected from the GamePhysics section on the Reddit website. Our approach shows promising results in our extensive analysis of simple queries, compound queries, and bug queries, indicating that our approach is useful for object and event detection in gameplay videos. An example application of our approach is as a gameplay video search engine to aid in reproducing video game bugs. Please visit the following link for the code and the data: https://asgaardlab.github.io/CLIPxGamePhysics/
Mohammad Reza Taesiri, Finlay Macklon, Cor-Paul Bezemer
MSR3
2022 Studying the Performance Risks of Upgrading Docker Hub Images: A Case Study of WordPress
abstract
The Docker Hub repository contains Docker images of applications, which allow users to do in-place upgrades to benefit from the latest released features and security patches. However, prior work showed that upgrading a Docker image not only changes the main application, but can also change many dependencies. In this paper, we present a methodology to study the performance impact of upgrading the Docker Hub image of an application, thereby focusing on changes to dependencies. We demonstrate our methodology through a case study of 90 official images of the WordPress application. Our study shows that Docker image users should be cautious and conduct a performance test before upgrading to a newer Docker image in most cases. Our methodology can assist them to better understand the performance risks of such upgrades, and helps them to decide how thorough such a performance test should be.
Mikael Sabuhi, Petr Musilek, Cor-Paul Bezemer
ICPE3
2022 A case study on the stability of performance tests for serverless applications
Simon Eismann, Diego Costa 0001, Lizhi Liao, Cor-Paul Bezemer, Weiyi Shang, André van Hoorn, Samuel Kounev
J. Syst. Softw.4
2022 What Causes Wrong Sentiment Classifications of Game Reviews?
abstract
Sentiment analysis is a popular technique to identify the sentiment of a piece of text. Several different domains have been targeted by sentiment analysis research, such as Twitter, movie reviews, and mobile app reviews. Although several techniques have been proposed, the performance of current sentiment analysis techniques is still far from acceptable, mainly when applied in domains on which they were not trained. In addition, the causes of wrong classifications are not clear. In this article, we study how sentiment analysis performs on game reviews. We first report the results of a large-scale empirical study on the performance of widely used sentiment classifiers on game reviews. Then, we investigate the root causes for the wrong classifications and quantify the impact of each cause on the overall performance. We study three existing classifiers:Stanford CoreNLP,NLTK, andSentiStrength. Our results show that most classifiers do not perform well on game reviews, with the best one beingNLTK(with an AUC of 0.70). We also identified four main causes for wrong classifications, such as reviews that point out advantages and disadvantages of the game, which might confuse the classifier. The identified causes are not trivial to be resolved and we call upon sentiment analysis and game researchers and developers to prioritize a research agenda that investigates how the performance of sentiment analysis of game reviews can be improved, for instance by developing techniques that can automatically deal with specific game-related issues of reviews (e.g., reviews with advantages and disadvantages). Finally, we show that training sentiment classifiers on reviews that are stratified by the game genre is effective.
Markos Viggiato, Dayi Lin, Abram Hindle, Cor-Paul Bezemer
IEEE Trans. Games4
2021 Improving the Discoverability of Indie Games by Leveraging their Similarity to Top-Selling Games: Identifying Important Requirements of a Recommender System
abstract
Indie games often lack visibility as compared to top-selling games due to their limited marketing budget and the fact that there are a large number of indie games. Players of top-selling games usually like certain types of games or certain game elements such as theme, gameplay, storyline. Therefore, indie games could leverage their shared game elements with top-selling games to get discovered. In this paper, we propose an approach to improve the discoverability of indie games by recommending similar indie games to gamers of top-selling games. We first matched 2,830 itch.io indie games to 326 top-selling Steam games. We then contacted the indie game developers for evaluation feedback and suggestions. We found that the majority of them (67.9%) who offered verbose responses show positive support for our approach. We also analyzed the reasons for bad recommendations and the suggestions by indie game developers to lay out the important requirements for such a recommendation system. The most important ones are: a standardized and extensive tag and genre ontology system is needed to bridge the two platforms, the expectations of players of top-selling games should be managed to avoid disappointment, a player’s preferences should be integrated when making recommendations, a standardized age restriction rule is needed, and finally, the recommendation tool should also show indie games that are the least similar or less popular.
Ngoc Quang Vu, Cor-Paul Bezemer
FDG2
2021 PySStuBs: Characterizing Single-Statement Bugs in Popular Open-Source Python Projects
abstract
Single-statement bugs (SStuBs) can have a severe impact on developer productivity. Despite usually being simple and not offering much of a challenge to fix, these bugs may still disturb a developer's workflow and waste precious development time. However, few studies have paid attention to these simple bugs, focusing instead on bugs of any size and complexity. In this study, we explore the occurrence of SStuBs in some of the most popular open-source Python projects on GitHub, while also characterizing their patterns and distribution. We further compare these bugs to SStuBs found in a previous study on Java Maven projects. We find that these Python projects have different SStuB patterns than the ones in Java Maven projects and identify 7 new SStuB patterns. Our results may help uncover the importance of understanding these bugs for the Python programming language, and how developers can handle them more effectively.
Arthur V. Kamienski, Luisa Palechor, Cor-Paul Bezemer, Abram Hindle
MSR3
2021 An empirical study of same-day releases of popular packages in the npm ecosystem
Filipe Roseiro Côgo, Gustavo Ansaldi Oliva, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2021 An empirical study of Q&A websites for game developers
Arthur V. Kamienski, Cor-Paul Bezemer
Empir. Softw. Eng.2
2021 How are issue reports discussed in Gitter chat rooms?
Hareem Sahar, Abram Hindle, Cor-Paul Bezemer
J. Syst. Softw.3
2021 An Empirical Study of Trends of Popular Virtual Reality Games and Their Complaints
abstract
The market for virtual reality (VR) games is growing rapidly and is expected to grow from 3.3 billion in 2018 to 13.7 billion in 2022. Due to the immersive nature of such games and the use of VR headsets, players may have complaints about VR games, which are distinct from those about traditional computer games, and an understanding of those complaints could enable developers to better take advantage of the growing VR market. We conduct an empirical study of 750 popular VR games and 17 635 user reviews on Steam in order to understand trends in VR games and their complaints. We find that the VR games market is maturing. Fewer VR games are released each month, but their quality appears to be improving over time. Most games support multiple headsets and play areas, and support for smaller scale play areas is increasing. Complaints of cybersickness are rare and declining, indicating that players are generally more concerned with other issues. Recently, complaints about game-specific issues have become the most frequent type of complaint, and VR game developers can now focus on these issues and worry less about VR-comfort issues such as cybersickness.
Rain Epp, Dayi Lin, Cor-Paul Bezemer
IEEE Trans. Games3
2021 What's Wrong with My Benchmark Results? Studying Bad Practices in JMH Benchmarks
abstract
Microbenchmarking frameworks, such as Java's Microbenchmark Harness (JMH), allow developers to write fine-grained performance test suites at the method or statement level. However, due to the complexities of the Java Virtual Machine, developers often struggle with writing expressive JMH benchmarks which accurately represent the performance of such methods or statements. In this paper, we empirically study bad practices of JMH benchmarks. We present a tool that leverages static analysis to identify 5 bad JMH practices. Our empirical study of 123 open source Java-based systems shows that each of these 5 bad practices are prevalent in open source software. Further, we conduct several experiments to quantify the impact of each bad practice in multiple case studies, and find that bad practices often significantly impact the benchmark results. To validate our experimental results, we constructed seven patches that fix the identified bad practices for six of the studied open source projects, of which six were merged into the main branch of the project. In this paper, we show that developers struggle with accurate Java microbenchmarking, and provide several recommendations to developers of microbenchmarking frameworks on how to improve future versions of their framework.
Diego Costa 0001, Cor-Paul Bezemer, Philipp Leitner 0001, Artur Andrzejak 0001
IEEE Trans. Software Eng.2
2021 Studying the Association Between Bountysource Bounties and the Issue-Addressing Likelihood of GitHub Issue Reports
abstract
Due to the voluntary nature of open source software, it can be hard to find a developer to work on a particular task. For example, some issue reports may be too cumbersome and unexciting for someone to volunteer to do them, yet these issue reports may be of high priority to the success of a project. To provide an incentive for implementing such issue reports, one can propose a monetary reward, i.e., a bounty, to the developer who completes that particular task. In this paper, we study bounties in open source projects on GitHub to better understand how bounties can be leveraged to evolve such projects in terms of addressing issue reports. We investigated 5,445 bounties for GitHub projects. These bounties were proposed through the Bountysource platform with a total bounty value of $406,425. We find that 1) in general, the timing of proposing bounties is the most important factor that is associated with the likelihood of an issue being addressed. More specifically, issue reports are more likely to be addressed if they are for projects in which bounties are used more frequently and if they are proposed earlier. 2) The bounty value of an issue report is the most important factor that is associated with the issue-addressing likelihood in the projects in which no bounties were used before. 3) There is a risk of wasting money for backers who invest money on long-standing issue reports.
Jiayuan Zhou, Shaowei Wang 0002, Cor-Paul Bezemer, Ying Zou 0001, Ahmed E. Hassan
IEEE Trans. Software Eng.3
2020 An Empirical Study of the Characteristics of Popular Game Jams and Their High-ranking Submissions on itch.io
abstract
Game jams are hackathon-like events that allow participants to develop a playable game prototype within a time limit. They foster creativity and the exchange of ideas by letting developers with different skill sets collaborate. Having a high-ranking game is a great bonus to a beginning game developer’s résumé and their pursuit of a career in the game industry. However, participants often face time constraints set by jam hosts while balancing what aspects of their games should be emphasized to have the highest chance of winning. Similarly, hosts need to understand what to emphasize when organizing online jams so that their jams are more popular, in terms of submission rate. In this paper, we study 1,290 past game jams and their 3,752 submissions on itch.io to understand better what makes popular jams and high-ranking games perceived well by the audience. We find that a quality description has a positive contribution to both a jam’s popularity and a game’s ranking. Additionally, more manpower organizing a jam or developing a game increases a jam’s popularity and a game’s high-ranking likelihood. High-ranking games tend to support Windows or macOS, and belong to the “Puzzle”, “Platformer”, “Interactive Fiction”, or “Action” genres. Also, shorter competitive jams tend to be more popular. Based on our findings, we suggest jam hosts and participants improve the description of their products and consider co-organizing or co-participating in a jam. Furthermore, jam participants should develop multi-platform multi-genre games. Finally, jam hosts should introduce a tighter time limit to increase their jam’s popularity.
Quang N. Vu, Cor-Paul Bezemer
FDG2
2020 Microservices: A Performance Tester's Dream or Nightmare?
abstract
In recent years, there has been a shift in software development towards microservice-based architectures, which consist of small services that focus on one particular functionality. Many companies are migrating their applications to such architectures to reap the benefits of microservices, such as increased flexibility, scalability and a smaller granularity of the offered functionality by a service. On the one hand, the benefits of microservices for functional testing are often praised, as the focus on one functionality and their smaller granularity allow for more targeted and more convenient testing. On the other hand, using microservices has their consequences (both positive and negative) on other types of testing, such as performance testing. Performance testing is traditionally done by establishing the baseline performance of a software version, which is then used to compare the performance testing results of later software versions. However, as we show in this paper, establishing such a baseline performance is challenging in microservice applications. In this paper, we discuss the benefits and challenges of microservices from a performance tester's point of view. Through a series of experiments on the TeaStore application, we demonstrate how microservices affect the performance testing process, and we demonstrate that it is not straightforward to achieve reliable performance testing results for a microservice application.
Simon Eismann, Cor-Paul Bezemer, Weiyi Shang, Dusan Okanovic, André van Hoorn
ICPE2
2020 A Framework for Satisfying the Performance Requirements of Containerized Software Systems Through Multi-Versioning
abstract
With the increasing popularity and complexity of containerized software systems, satisfying the performance requirements of these systems becomes more challenging as well. While a common remedy to this problem is to increase the allocated amount of resources by scaling up or out, this remedy is not necessarily cost-effective and, therefore, often problematic for smaller companies. In this paper, we study an alternative, more cost-effective approach for satisfying the performance requirements of containerized software systems. In particular, we investigate how we can satisfy such requirements by applying software multi-versioning to the system's resource-heavy containers. We present DockerMV, an open-source extension of the Docker framework, to support the multi-versioning of containerized software systems. We demonstrate the efficacy of multi-versioning for satisfying the performance requirements of containerized software systems through experiments on the TeaStore, a microservice reference test application, and Znn, a containerized news portal application. Our DockerMV extension can be used by software developers to introduce multi-versioning in their own containerized software systems, thereby better allowing them to meet the performance requirements of their systems.
Sara Gholami, Alireza Goli, Cor-Paul Bezemer, Hamzeh Khazaei
ICPE3
2020 A longitudinal study of popular ad libraries in the Google Play Store
Md. Ahasanuzzaman, Safwat Hassan, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2020 Building the perfect game - an empirical study of game modifications
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2020 An empirical study of the characteristics of popular Minecraft mods
Gopi Krishnan Rajbahadur, Dayi Lin, Mohammed Sayagh, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.5
2020 Bounties on technical Q&A sites: a case study of Stack Overflow bounties
Jiayuan Zhou, Shaowei Wang 0002, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2020 Studying Bad Updates of Top Free-to-Download Apps in the Google Play Store
abstract
Developers always focus on delivering high-quality updates to improve, or maintain the rating of their apps. Prior work has studied user reviews by analyzing all reviews of an app. However, this app-level analysis misses the point that users post reviews to provide their feedback on a certain update. For example, two bad updates of an app with a history of good updates would not be spotted using app-level analysis. In this paper, we examine reviews at the update-level to better understand how users perceive bad updates. We focus our study on the top 250 bad updates (i.e., updates with the highest increase in the percentage of negative reviews relative to the prior updates of the app) from 26,726 updates of 2,526 top free-to-download apps in the Google Play Store. We find that feature removal and UI issues have the highest increase in the percentage of negative reviews. Bad updates with crashes and functional issues are the most likely to be fixed by a later update. However, developers often do not mention these fixes in the release notes. Our work demonstrates the necessity of an update-level analysis of reviews to capture the feelings of an app's user-base about a particular update.
Safwat Hassan, Cor-Paul Bezemer, Ahmed E. Hassan
IEEE Trans. Software Eng.2
2019 How is Performance Addressed in DevOps?
abstract
DevOps is a modern software engineering paradigm that is gaining widespread adoption in industry. The goal of DevOps is to bring software changes into production with a high frequency and fast feedback cycles. This conflicts with software quality assurance activities, particularly with respect to performance. For instance, performance evaluation activities --- such as load testing --- require a considerable amount of time to get statistically significant results.
Cor-Paul Bezemer, Simon Eismann, Vincenzo Ferme, Johannes Grohmann, Robert Heinrich, Pooyan Jamshidi, Weiyi Shang, André van Hoorn, Mónica Villavicencio, Jürgen Walter, Felix Willnecker
ICPE1
2019 Studying the consistency of star ratings and reviews of popular free hybrid Android and iOS apps
Hanyang Hu, Shaowei Wang 0002, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2019 The impact of feature reduction techniques on defect prediction models
Masanari Kondo, Cor-Paul Bezemer, Yasutaka Kamei, Ahmed E. Hassan, Osamu Mizuno
Empir. Softw. Eng.2
2019 Identifying gameplay videos that exhibit bugs in computer games
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.2
2019 An empirical study of game reviews on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ying Zou 0001, Ahmed E. Hassan
Empir. Softw. Eng.2
2019 How do developers utilize source code from stack overflow?
abstract
Technical question and answer Q&A platforms, such as Stack Overflow, provide a platform for users to ask and answer questions about a wide variety of programming topics. These platforms accumulate a large amount of knowledge, including hundreds of thousands lines of source code. Developers can benefit from the source code that is attached to the questions and answers on Q&A platforms by copying or learning from (parts of) it. By understanding how developers utilize source code from Q&A platforms, we can provide insights for researchers which can be used to improve next-generation Q&A platforms to help developers reuse source code fast and easily. In this paper, we first conduct an exploratory study on 289 files from 182 open-source projects, which contain source code that has an explicit reference to a Stack Overflow post. Our goal is to understand how developers utilize code from Q&A platforms and to reveal barriers that may make code reuse more difficult. In 31.5% of the studied files, developers needed to modify source code from Stack Overflow to make it work in their own projects. The degree of required modification varied from simply renaming variables to rewriting the whole algorithm. Developers sometimes chose to implement an algorithm from scratch based on the descriptions from Stack Overflow answers, even if there was an implementation readily available in the post. In 35.5% of the studied files, developers used Stack Overflow posts as an information source for later reference. To further understand the barriers of reusing code and to obtain suggestions for improving the code reuse process on Q&A platforms, we conducted a survey with 453 open-source developers who are also on Stack Overflow. We found that the top 3 barriers that make it difficult for developers to reuse code from Stack Overflow are: (1) too much code modification required to fit in their projects, (2) incomprehensive code, and (3) low code quality. We summarized and analyzed all survey responses and we identified that developers suggest improvements for future Q&A platforms along the following dimensions: code quality, information enhancement & management, data organization, license, and the human factor. For instance, developers suggest to improve the code quality by adding an integrated validator that can test source code online, and an outdated code detection mechanism. Our findings can be used as a roadmap for researchers and developers to improve code reuse.
Shaowei Wang 0002, Cor-Paul Bezemer, Katsuro Inoue
Empir. Softw. Eng.3
2018 Analyzing a decade of Linux system calls
abstract
The Linux kernel provides its services to the application layer using so-called system calls. All system calls combined form the Application Programming Interface (API) of the kernel. Hence, system calls provide us with a window into the development process and design decisions that are made for the Linux kernel. Our paper [1] presents the result of an empirical study of the changes (8,770) that were made to the system calls during the last decade (i.e., from April 2005 to December 2014). The main contributions and most important findings of our study are:
Mojtaba Bagherzadeh, Nafiseh Kahani, Cor-Paul Bezemer, Ahmed E. Hassan, Jürgen Dingel, James R. Cordy
ICSE3
2018 Studying the dialogue between users and developers of free apps in the google play store
abstract
The popularity of mobile apps continues to grow over the past few years. Mobile app stores, such as the Google Play Store and Apple's App Store provide a unique user feedback mechanism to app developers through app reviews. In the Google Play Store (and most recently in the Apple App Store), developers are able to respond to such user feedback.
Safwat Hassan, Chakkrit Tantithamthavorn, Cor-Paul Bezemer, Ahmed E. Hassan
ICSE3
2018 An empirical study of early access games on the steam platform
abstract
"Early access" is a release strategy for software that allows consumers to purchase an unfinished version of the software. In turn, consumers can influence the software development process by giving developers early feedback. This early access model has become increasingly popular through digital distribution platforms, such as Steam which is the most popular distribution platform for games. The plethora of options offered by Steam to communicate between developers and game players contribute to the popularity of the early access model.
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan
ICSE2
2018 Analyzing a decade of Linux system calls
Mojtaba Bagherzadeh, Nafiseh Kahani, Cor-Paul Bezemer, Ahmed E. Hassan, Jürgen Dingel, James R. Cordy
Empir. Softw. Eng.3
2018 Studying the dialogue between users and developers of free apps in the Google Play Store
Safwat Hassan, Chakkrit Tantithamthavorn, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.3
2018 Studying the consistency of star ratings and the complaints in 1 & 2-star user reviews for top free cross-platform Android and iOS apps
Hanyang Hu, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.2
2018 Examining the stability of logging statements
Suhas Kabinna, Cor-Paul Bezemer, Weiyi Shang, Mark D. Syer, Ahmed E. Hassan
Empir. Softw. Eng.2
2018 An empirical study of early access games on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.2
2018 Revisiting the performance of automated approaches for the retrieval of duplicate reports in issue tracking systems that perform just-in-time duplicate retrieval
Mohamed Sami Rakha, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.2
2018 Revisiting the Performance Evaluation of Automated Approaches for the Retrieval of Duplicate Issue Reports
abstract
Issue tracking systems (ITSs), such as Bugzilla, are commonly used to track reported bugs, improvements and change requests for a software project. To avoid wasting developer resources on previously-reported (i.e., duplicate) issues, it is necessary to identify such duplicates as soon as they are reported. Several automated approaches have been proposed for retrieving duplicate reports, i.e., identifying the duplicate of a new issue report in a list of$n$candidates. These approaches rely on leveraging the textual, categorical, and contextual information in previously-reported issues to decide whether a newly-reported issue has previously been reported. In general, these approaches are evaluated using data that spans a relatively short period of time (i.e., the classical evaluation). However, in this paper, we show that the classical evaluation tends to overestimate the performance of automated approaches for retrieving duplicate issue reports. Instead, we propose a realistic evaluation using all the reports that are available in the ITS of a software project. We conduct experiments in which we evaluate two popular approaches for retrieving duplicate issues (BM25F and REP) using the classical and realistic evaluations. We find that for the issue tracking data of the Mozilla foundation, the Eclipse foundation and OpenOffice, the realistic evaluation shows that previously proposed approaches perform considerably lower than previously reported using the classical evaluation. As a result, we conclude that the reported performance of approaches for retrieving duplicate issue reports is significantly overestimated in literature. In order to improve the performance of the automated retrieval of duplicate issue reports, we propose to leverage the resolution field of issue reports. Our experiments show that a relative improvement in the performance of a median of 7-21.5 percent and a maximum of 19-60 percent can be achieved by leveraging the resolution field of issue reports for the automated retrieval of duplicates.
Mohamed Sami Rakha, Cor-Paul Bezemer, Ahmed E. Hassan
IEEE Trans. Software Eng.2
2017 An Exploratory Study of the State of Practice of Performance Testing in Java-Based Open Source Projects
abstract
The usage of open source (OS) software is wide-spread across many industries. While the functional quality of OS projects is considered to be similar to closed-source software, much is unknown about the quality in terms of performance. One challenge for OS developers is that, unlike for functional testing, there is a lack of accepted best practices for performance testing. To reveal the state of practice of performance testing in OS projects, we conduct an exploratory study on 111 Java-based OS projects from GitHub. We study the performance tests of these projects from five perspectives: (1) developers, (2) size, (3) test organization, (4) types of performance tests and (5) used tooling. We show that writing performance tests is not a popular task in OS projects: performance tests form only a small portion of the test suite, are rarely updated, and are usually maintained by a small group of core project developers. Further, even though many projects are aware that they need performance tests, developers appear to struggle implementing them. We argue that future performance testing frameworks should provider better support for low-friction testing, for instance via non-parameterized methods or performance test generation, as well as focus on a tight integration with standard continuous integration tooling.
Philipp Leitner 0001, Cor-Paul Bezemer
ICPE2
2017 An empirical study of unspecified dependencies in make-based build systems
Cor-Paul Bezemer, Shane McIntosh, Bram Adams, Daniel M. Germán, Ahmed E. Hassan
Empir. Softw. Eng.1
2017 Studying the urgent updates of popular games on the Steam platform
Dayi Lin, Cor-Paul Bezemer, Ahmed E. Hassan
Empir. Softw. Eng.2
2016 Studying the effectiveness of application performance management (APM) tools for detecting performance regressions for web applications: an experience report
abstract
Performance regressions, such as a higher CPU utilization than in the previous version of an application, are caused by software application updates that negatively affect the performance of an application. Although a plethora of mining software repository research has been done to detect such regressions, research tools are generally not readily available to practitioners. Application Performance Management (APM) tools are commonly used in practice for detecting performance issues in the field by mining operational data.
Tarek M. Ahmed, Cor-Paul Bezemer, Tse-Hsun (Peter) Chen, Ahmed E. Hassan, Weiyi Shang
MSR2
2016 Logging library migrations: a case study for the apache software foundation projects
abstract
Developers leverage logs for debugging, performance monitoring and load testing. The increased dependence on logs has lead to the development of numerous logging libraries which help developers in logging their code. As new libraries emerge and current ones evolve, projects often migrate from an older library to another one.
Suhas Kabinna, Cor-Paul Bezemer, Weiyi Shang, Ahmed E. Hassan
MSR2
2016 Examining the Stability of Logging Statements
abstract
Logging statements produce logs that assist in understanding system behavior, monitoring choke-points and debugging. Prior research demonstrated the importance of logging statements in operating, understanding and improving software systems. The importance of logs has lead to a new market of log management and processing tools. However, logs are often unstable, i.e., the logging statements that generate logs are often changed without the consideration of other stakeholders, causing misleading results and failures of log processing tools. In order to proactively mitigate such issues that are caused by unstable logging statements, in this paper we empirically study the stability of logging statements in four open source applications namely:Liferay, ActiveMQ, Camel and Cloud Stack. We find that 20-45% of the logging statements in our studied applications change throughout their lifetime. The median number of days between the introduction of a logging statement and the first change to that statement is between 1 and 17 in our studied applications. These numbers show that in order to reduce maintenance effort, developers of log processing tools must be careful when selecting the logging statements on which they will let their tools depend. In this paper, we make an important first step towards assisting developers of log processing tools in determining whether a logging statement is likely to remain unchanged in the future. Using random forest classifiers, we examine which metrics are important for understanding whether a logging statement will change. We show that our classifiers achieve 83%-91% precision and 65%-85% recall in the four studied applications. We find that file ownership, developer experience, log density and SLOC are important metrics for determining whether a logging statement will change in the future. Developers can use this knowledge to build more robust log processing tools, by making those tools depend on logs that are generated by logging statements that are likely to remain unchanged.
Suhas Kabinna, Weiyi Shang, Cor-Paul Bezemer, Ahmed E. Hassan
SANER3
2016 Optimizing the Performance-Related Configurations of Object-Relational Mapping Frameworks Using a Multi-Objective Genetic Algorithm
abstract
Object-relational mapping (ORM) frameworks map low-level database operations onto a high-level programming API that can be accessed from within object-oriented source code. ORM frameworks often provide configuration options to optimize the performance of such database operations. However, determining the set of optimal configuration options is a challenging task. Through an exploratory study on two open source applications (Spring PetClinic and ZK), we find that the difference in execution time between two configurations can be large. In addition, both applications are not shipped with an ORM configuration that is related to performance: instead, they use the default values provided by the ORM framework. We show that in 89% of the 9 analyzed test cases for PetClinic and in 96% of the 54 analyzed test cases for ZK, the default configuration values supplied by the ORM framework performed significantly slower than the optimal configuration for that test case. Based on these observations, this paper proposes an approach for automatically finding an optimal ORM configuration using a multi-objective genetic algorithm. We evaluate our approach by conducting a case study of Spring PetClinic and ZK. We find that our approach finds near-optimal configurations in 360-450 seconds for PetClinic and in 9-12 hours for ZK. These execution times allow our approach to be executed to find an optimal configuration before each new release of an application.
Ravjot Singh, Cor-Paul Bezemer, Weiyi Shang, Ahmed E. Hassan
ICPE2
2015 Understanding software performance regressions using differential flame graphs
abstract
Flame graphs are gaining rapidly in popularity in industry to visualize performance profiles collected by stack-trace based profilers. In some cases, for example, during performance regression detection, profiles of different software versions have to be compared. Doing this manually using two or more flame graphs or textual profiles is tedious and error-prone. In this `Early Research Achievements'-track paper, we present our preliminary results on using differential flame graphs instead. Differential flame graphs visualize the differences between two performance profiles. In addition, we discuss which research fields we expect to benefit from using differential flame graphs. We have implemented our approach in an open source prototype called FLAMEGRAPHDIFF, which is available on GitHub. FLAMEGRAPHDIFF makes it easy to generate interactive differential flame graphs from two existing performance profiles. These graphs facilitate easy tracing of elements in the different graphs to ease the understanding of the (d)evolution of the performance of an application.
Cor-Paul Bezemer, Johan A. Pouwelse, Brendan Gregg
SANER1
2015 Defining multi-tenancy: A systematic mapping study on the academic and the industrial perspective
Jaap Kabbedijk, Cor-Paul Bezemer, Slinger Jansen, Andy Zaidman
J. Syst. Softw.2
2014 Performance optimization of deployed software-as-a-service applications
Cor-Paul Bezemer, Andy Zaidman
J. Syst. Softw.1
2014 Detecting and analyzing I/O performance regressions
abstract
Regression testing can be done by re-executing a test suite on different software versions and comparing the outcome. For functional testing, the outcome of such tests is either pass (correct behaviour) or fail (incorrect behaviour). For non-functional testing, such as performance testing, this is more challenging as correct and incorrect are not clearly defined concepts for these types of testing. In this paper, we present an approach for detecting and analyzing I/O performance regressions. Our method is supplemental to existing profilers and its goal is to analyze the effect of source code changes on the performance of a system. In this paper, we focus on analyzing the amount of I/O writes being done. The open source implementation of our approach, SPECTRAPERF, is available for download. We evaluate our approach in a field user study on Tribler, an open source peer-to-peer client and its decentralized solution for synchronizing messages, Dispersy. In this evaluation, we show that our approach can guide the performance optimization process, as it helps developers to find performance bottlenecks on the one hand, and on the other allows them to validate the effect of performance optimizations. In addition, we perform a feasibility study on Django, the most popular Python project on Github, to demonstrate our applicability on other projects. Copyright c 2013 John Wiley & Sons, Ltd.
Cor-Paul Bezemer, Elric Milon, Andy Zaidman, Johan A. Pouwelse
J. Softw. Evol. Process.1
2012 Locating performance improvement opportunities in an industrial software-as-a-service application
abstract
The goal of performance maintenance is to improve the performance of a software system after delivery. As the performance of a system is often characterized by unexpected combinations of metric values, manual analysis of performance is hard in complex systems. In this paper, we extend our previous work on performance anomaly detection with a technique that helps performance experts locate spots - so-called performance improvement opportunities (PIOs) -, for possible performance improvements. PIOs give performance experts a starting point for performance improvements, e.g., by pinpointing the bottleneck component. The technique uses a combination of association rules and several visualizations, such as heat maps, which were implemented in an open source tool called Wedjat. In this paper, we evaluate our technique and Wedjat in a field user study with three performance experts from industry using data from a large-scale industrial application. From our field study we conclude that our technique is useful for speeding up the performance maintenance process and that heat maps are a valuable way of visualizing performance data.
Cor-Paul Bezemer, Andy Zaidman, Ad van der Hoeven, Andre van de Graaf, Maarten Wiertz, Remko Weijers
ICSM1
2010 Enabling multi-tenancy: An industrial experience report
abstract
Multi-tenancy is a relatively new software architecture principle in the realm of the Software as a Service (SaaS) business model. It allows to make full use of the economy of scale, as multiple customers - “tenants” - share the same application and database instance. All the while, the tenants enjoy a highly configurable application, making it appear that the application is deployed on a dedicated server. The major benefits of multi-tenancy are increased utilization of hardware resources and improved ease of maintenance, resulting in lower overall application costs, making the technology attractive for service providers targeting small and medium enterprises (SME). Therefore, migrating existing single-tenant to multi-tenant applications can be interesting for SaaS software companies. In this paper we report on our experiences with reengineering an existing industrial, single-tenant software system into a multitenant one using a lightweight reengineering approach.
Cor-Paul Bezemer, Andy Zaidman, Bart Platzbeecker, Toine Hurkmans, Aad 't Hart
ICSM1
2009 Automated security testing of web widget interactions
abstract
We present a technique for automatically detecting security vulnerabilities in client-side self-contained components, called web widgets, that can co-exist independently on a single web page. In this paper we focus on two security scenarios, namely the case in which (1) a malicious widget changes the content (DOM) of another widget, and (2) a widget steals data from another widget and sends it to the server via an HTTP request. We propose a dynamic analysis approach for automatically executing the web application and analyzing the runtime changes in the user interface, as well as the outgoing HTTP calls, to detect inter-widget interaction violations.
Cor-Paul Bezemer, Ali Mesbah 0001, Arie van Deursen
ESEC/SIGSOFT FSE1