VLDB 2026 Research / reviewers in the wild / expert
Markos Viggiato
dblp:223/6503
· DBLP profile ↗
11ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-8500-3723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Leveraging the OPT Large Language Model for Sentiment Analysis of Game ReviewsabstractAutomatically extracting players' sentiments about games can help game developers to better understand the aspects of their games that players like or dislike. Our prior work showed that traditional sentiment analysis techniques do not perform well on game reviews. However, the Natural Language Processing (NLP) field has seen a steep progress in recent years. In this letter, we follow up on our prior work and investigate how a state-of-the-art large language model (OPT-175B) performs on the sentiment classification of game reviews. We manually analyze the game reviews wrongly classified by OPT-175B to better understand the issues that affect the performance of that model and how those issues compare to the challenges faced by traditional classifiers. We found that OPT-175B achieves (far) better performance than traditional sentiment classifiers, with a 72%-increased F-measure and a 30%-increased AUC compared to the best traditional classifier studied in our prior work. We also found that common challenges of traditional classifiers, such as reviews with game comparisons and negative terminology, have been mostly solved by the OPT-175B model. Markos Viggiato, Cor-Paul Bezemer |
IEEE Trans. Games | 1 |
| 2023 | Prioritizing Natural Language Test Cases Based on Highly-Used Game FeaturesabstractSoftware testing is still a manual activity in many industries, such as the gaming industry. But manually executing tests becomes impractical as the system grows and resources are restricted, mainly in a scenario with short release cycles. Test case prioritization is a commonly used technique to optimize the test execution. However, most prioritization approaches do not work for manual test cases as they require source code information or test execution history, which is often not available in a manual testing scenario. In this paper, we propose a prioritization approach for manual test cases written in natural language based on the tested application features (in particular, highly-used application features). Our approach consists of (1) identifying the tested features from natural language test cases (with zero-shot classification techniques) and (2) prioritizing test cases based on the features that they test. We leveraged the NSGA-II genetic algorithm for the multi-objective optimization of the test case ordering to maximize the coverage of highly-used features while minimizing the cumulative execution time. Our findings show that we can successfully identify the application features covered by test cases using an ensemble of pre-trained models with strong zero-shot capabilities (an F-score of 76.1%). Also, our prioritization approaches can find test case orderings that cover highly-used application features early in the test execution while keeping the time required to execute test cases short. QA engineers can use our approach to focus the test execution on test cases that cover features that are relevant to users. Markos Viggiato, Dale Paas, Cor-Paul Bezemer |
ESEC/SIGSOFT FSE | 1 |
| 2023 | A Taxonomy of Testable HTML5 Canvas IssuesabstractThe HTML5canvas> is widely used to display high quality graphics in web applications. However, the combination of web, GUI, and visual techniques that are required to buildcanvas> applications, together with the lack of testing and debugging tools, makes developing such applications very challenging. To help direct future research on testingcanvas> applications, in this paper we present a taxonomy of testablecanvas> issues. First, we extracted 2,403canvas>-related issue reports from 123 open source GitHub projects that use the HTML5canvas>. Second, we constructed our taxonomy by manually classifying a random sample of 332 issue reports. Our manual classification identified five broad categories of testablecanvas> issues, such as Visual and Performance issues. We found that Visual issues are the most frequent (35%), while Performance issues are relatively infrequent (5%). We also found that many testablecanvas> issues that present themselves visually on thecanvas> are actually caused by other components of the web application. Our taxonomy of testablecanvas> issues can be used to steer future research intocanvas> issues and testing. Finlay Macklon, Markos Viggiato, Natalia Romanova, Chris Buzon, Dale Paas, Cor-Paul Bezemer |
IEEE Trans. Software Eng. | 2 |
| 2023 | Identifying Similar Test Cases That Are Specified in Natural LanguageabstractSoftware testing is still a manual process in many industries, despite the recent improvements in automated testing techniques. As a result, test cases (which consist of one or more test steps that need to be executed manually by the tester) are often specified in natural language by different employees and many redundant test cases might exist in the test suite. This increases the (already high) cost of test execution. Manually identifying similar test cases is a time-consuming and error-prone task. Therefore, in this paper, we propose an unsupervised approach to identify similar test cases. Our approach uses a combination of text embedding, text similarity and clustering techniques to identify similar test cases. We evaluate five different text embedding techniques, two text similarity metrics, and two clustering techniques to cluster similar test steps and three techniques to identify similar test cases from the test step clusters. Through an evaluation in an industrial setting, we showed that our approach achieves a high performance to cluster test steps (an F-score of 87.39%) and identify similar test cases (an F-score of 86.13%). Furthermore, a validation with developers indicates several different practical usages of our approach (such as identifying redundant test cases), which help to reduce the testing manual effort and time. Markos Viggiato, Dale Paas, Chris Buzon, Cor-Paul Bezemer |
IEEE Trans. Software Eng. | 1 |
| 2022 | Automatically Detecting Visual Bugs in HTML5 Canvas GamesabstractThe HTML5 is used to display high quality graphics in web applications such as web games (i.e., games). However, automatically testing games is not possible with existing web testing techniques and tools, and manual testing is laborious. Many widely used web testing tools rely on the Document Object Model (DOM) to drive web test automation, but the contents of the are not represented in the DOM. The main alternative approach, snapshot testing, involves comparing oracle snapshot images with test-time snapshot images using an image similarity metric to catch visual bugs, i.e., bugs in the graphics of the web application. However, creating and maintaining oracle snapshot images for games is onerous, defeating the purpose of test automation. In this paper, we present a novel approach to automatically detect visual bugs in games. By leveraging an internal representation of objects on the , we decompose snapshot images into a set of object images, each of which is compared with a respective oracle asset (e.g., a sprite) using four similarity metrics: percentage overlap, mean squared error, structural similarity, and embedding similarity. We evaluate our approach by injecting 24 visual bugs into a custom game, and find that our approach achieves an accuracy of 100%, compared to an accuracy of 44.6% with traditional snapshot testing. Finlay Macklon, Mohammad Reza Taesiri, Markos Viggiato, Stefan Antoszko, Natalia Romanova, Dale Paas, Cor-Paul Bezemer |
ASE | 3 |
| 2022 | What Causes Wrong Sentiment Classifications of Game Reviews?abstractSentiment analysis is a popular technique to identify the sentiment of a piece of text. Several different domains have been targeted by sentiment analysis research, such as Twitter, movie reviews, and mobile app reviews. Although several techniques have been proposed, the performance of current sentiment analysis techniques is still far from acceptable, mainly when applied in domains on which they were not trained. In addition, the causes of wrong classifications are not clear. In this article, we study how sentiment analysis performs on game reviews. We first report the results of a large-scale empirical study on the performance of widely used sentiment classifiers on game reviews. Then, we investigate the root causes for the wrong classifications and quantify the impact of each cause on the overall performance. We study three existing classifiers:Stanford CoreNLP,NLTK, andSentiStrength. Our results show that most classifiers do not perform well on game reviews, with the best one beingNLTK(with an AUC of 0.70). We also identified four main causes for wrong classifications, such as reviews that point out advantages and disadvantages of the game, which might confuse the classifier. The identified causes are not trivial to be resolved and we call upon sentiment analysis and game researchers and developers to prioritize a research agenda that investigates how the performance of sentiment analysis of game reviews can be improved, for instance by developing techniques that can automatically deal with specific game-related issues of reviews (e.g., reviews with advantages and disadvantages). Finally, we show that training sentiment classifiers on reviews that are stratified by the game genre is effective. Markos Viggiato, Dayi Lin, Abram Hindle, Cor-Paul Bezemer |
IEEE Trans. Games | 1 |
| 2020 | Understanding machine learning software defect predictions
Geanderson E. dos Santos, Eduardo Figueiredo 0001, Adriano Veloso, Markos Viggiato, Nivio Ziviani |
Autom. Softw. Eng. | 4 |
| 2019 | Understanding similarities and differences in software development practices across domainsabstractSince software engineering is globalized and not a homogeneous whole, we expect that development practices are differently adopted across domains. However, little is known about how practices are followed in different software domains (e.g., healthcare, banking, and Oil and gas). In this paper, we report the results of an exploratory and inductive research, in which we seek differences and similarities regarding the adoption of several widespread practices across 13 domains. We interviewed 19 worldwide developers with experience in multiple domains (i.e., cross-domain developers) from large multinational companies, such as Facebook, Google, and Macy's. We also run a Web survey to confirm (or not) the interview results. Our findings show that, in fact, different domains adopt practices in a different fashion. We identified that continuous integration practices are interrupted during important commerce periods (e.g., Black Friday) in the financial domains. We also noticed the company's culture and policies strongly influence the adopted practices, instead of the domain itself. Our study also has important implications for global software engineering practices. For instance, companies should provide targeted training for their development teams and new interdisciplinary courses in software engineering and other domains, such as healthcare, are highly recommended. Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo 0001, Pooyan Jamshidi, Christian Kästner |
ICGSE | 1 |
| 2019 | How Do Code Changes Evolve in Different Platforms? A Mining-Based InvestigationabstractCode changes are performed differently in the mobile and non-mobile platforms. Prior work has investigated the differences in specific platforms. However, we still lack a deeper understanding of how code changes evolve across different software platforms. In this paper, we present a study aiming at investigating the frequency of changes and how source code, build and test changes co-evolve in mobile and non-mobile platforms. We developed regression models to explain which factors influence the frequency of changes and applied the Apriori algorithm to find types of changes that frequently co-occur. Our findings show that non-mobile repositories have a higher number of commits per month and our regression models suggest that being mobile significantly impacts on the number of commits in a negative direction when controlling for confound factors, such as code size. We also found that developers do not usually change source code files together with build or test files. We argue that our results can provide valuable information for developers on how changes are performed in different platforms so that practices adopted in successful software systems can be followed. Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo 0001, Pooyan Jamshidi, Christian Kästner |
ICSME | 1 |
| 2018 | Evaluating domain-specific metric thresholds: an empirical studyabstractSoftware metrics and thresholds provide means to quantify several quality attributes of software systems. Indeed, they have been used in a wide variety of methods and tools for detecting different sorts of technical debts, such as code smells. Unfortunately, these methods and tools do not take into account characteristics of software domains, as the intrinsic complexity of geo-localization and scientific software systems or the simple protocols employed by messaging applications. Instead, they rely on generic thresholds that are derived from heterogeneous systems. Although derivation of reliable thresholds has long been a concern, we still lack empirical evidence about threshold variation across distinct software domains. To tackle this limitation, this paper investigates whether and how thresholds vary across domains by presenting a large-scale study on 3,107 software systems from 15 domains. We analyzed the derivation and distribution of thresholds based on 8 well-known source code metrics. As a result, we observed that software domain and size are relevant factors to be considered when building benchmarks for threshold derivation. Moreover, we also observed that domain-specific metric thresholds are more appropriated than generic ones for code smell detection. Allan Mori, Gustavo Vale, Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo 0001, Elder Cirilo, Pooyan Jamshidi, Christian Kästner |
TechDebt@ICSE | 3 |
| 2018 | An Empirical Study on the Impact of Android Code Smells on Resource UsageabstractCode smells are symptoms that something may be wrong with the app.Aiming at removing code smells and improving the maintainability and performance of the app, we may apply the refactoring technique, which could reduce hardware resource use, such as CPU and memory.However, a few studies have evaluated the impacts of the refactoring in Android.This paper presents a study to assess the effects of smartphone resource use caused by refactoring of 3 classic code smells: God Class, God Method, and Feature Envy.To this purpose, we selected 9 apps from GitHub.The results show that refactoring used in desktop software may not be appropriate for Android apps.For example, the refactoring of God Method had increased CPU consumption by more than 47%, while the refactoring of the 3 code smells reduced memory consumption in average 6.51%, 8.4%, and 6.37%, respectively, in one app.Our results can support the community in conducting research and future implementation of new tools.Also, it guides app developers in refactoring and thus improving the quality of their apps. Johnatan Oliveira, Markos Viggiato, Mateus F. Santos, Eduardo Figueiredo 0001, Humberto Torres Marques-Neto |
SEKE | 2 |