VLDB 2026 Research / reviewers in the wild / expert
Gabor Antal
dblp:156/6366 · also Gábor Antal
· DBLP profile ↗
14ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-3002-8624ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary StudyabstractRecent advancements in large language models (LLMs) have shown promise for automated vulnerability detection and repair in software systems. This paper investigates the performance of GPT-4o in repairing Java vulnerabilities from a widely used dataset (Vul4J), exploring how different contextual information affects automated vulnerability repair (AVR) capabilities. We compare the latest GPT-4o’s performance against previous results with GPT-4 using identical prompts. We evaluated nine additional prompts crafted by us that contain various contextual information such as CWE or CVE information, and manually extracted code contexts. Each prompt was executed three times on 42 vulnerabilities, and the resulting fix candidates were validated using Vul4J’s automated testing framework. Our results show that GPT-4o performed 11.9% worse on average than GPT-4 with the same prompt, but was able to fix 10.5% more distinct vulnerabilities in the three runs together. CVE information significantly improved repair rates, while the length of the task description had minimal impact. Combining CVE guidance with manually extracted code context resulted in the best performance. Using our Top-3 prompts together, GPT-4o repaired 26 (62%) vulnerabilities at least once, outperforming both the original baseline (40%) and its reproduction (45%), suggesting that ensemble prompt strategies could improve vulnerability repair in zero-shot settings. Gabor Antal, Bence Bogenfürst, Rudolf Ferenc, Péter Hegedüs |
EASE | 1 |
| 2025 | Leveraging GPT-4 for Vulnerability-Witnessing Unit Test GenerationabstractIn the life-cycle of software development, testing plays a crucial role in quality assurance. Proper testing not only increases code coverage and prevents regressions but it can also ensure that any potential vulnerabilities in the software are identified and effectively fixed. However, creating such tests is a complex, resource-consuming manual process. To help developers and security experts, this paper explores the automatic unit test generation capability of one of the most widely used large language models, GPT-4, from the perspective of vulnerabilities. We examine a subset of the VUL4J dataset containing real vulnerabilities and their corresponding fixes to determine whether GPT-4 can generate syntactically and/or semantically correct unit tests based on the code before and after the fixes as evidence of vulnerability mitigation. We focus on the impact of code contexts, the effectiveness of GPT-4’s self-correction ability, and the subjective usability of the generated test cases. Our results indicate that GPT-4 can generate syntactically correct test cases 66.5% of the time without domain-specific pre-training. Although the semantic correctness of the fixes could be automatically validated in only 7. 5% of the cases, our subjective evaluation shows that GPT-4 generally produces test templates that can be further developed into fully functional vulnerability-witnessing tests with relatively minimal manual effort. Gabor Antal, Dénes Bán, Martin Isztin, Rudolf Ferenc, Péter Hegedüs |
EASE | 1 |
| 2024 | Reality Check: Assessing GPT-4 in Fixing Real-World Software VulnerabilitiesabstractDiscovering and mitigating software vulnerabilities is a challenging task. These vulnerabilities are often caused by simple, otherwise (and in other contexts) harmless code snippets (e.g., unchecked path traversal). Large Language Models (LLMs) promise to revolutionize not just human-machine interactions but various software engineering tasks as well, including the automatic repair of vulnerabilities. However, currently, it is hard to assess the performance, robustness, and reliability of these models as most of their evaluation has been done on small, synthetic examples. In our work, we systematically evaluate the automatic vulnerability fixing capabilities of GPT-4, a popular LLM, using a database of real-world Java vulnerabilities, Vul4J. We expect the model to provide fixes for vulnerable methods, which we evaluate manually and based on unit test results included in the Vul4J database. GPT-4 provided perfect fixes consistently for at least 12 out of the total 46 examined vulnerabilities, which could be applied as is. In an additional 5 cases, the provided textual instructions would help to fix the vulnerabilities in a practical scenario (despite the provided code being incorrect). Our findings, similar to others, also show that prompting has a significant effect. Zoltán Ságodi, Gabor Antal, Bence Bogenfürst, Martin Isztin, Péter Hegedüs, Rudolf Ferenc |
EASE | 2 |
| 2024 | On the Usefulness of Python Structural Pattern Matching: An Empirical StudyabstractAs the important role of software in our modern world becomes more and more evident, the need for more complex data structures is increasing. Structural pattern matching has become an elegant technique for simplifying complex conditionallogic in the source code in modern programming languages. It provides an easy and concise way to destructure complex data and enables making decisions based on its structure, ultimately improving code readability. In the context of Python, a language renowned for its simplicity, structural pattern matching was a missing feature until October 2021. Finally, the feature is shipped in Python 3.10, allowing developers to exploit the potential of structural pattern matching. Instead of traditional conditional branching and endless type checking, structural pattern matching offers a more straightforward way to handle data. In this paper, we investigate the usefulness of this relatively new Python feature by involving 65 participants in a code review experiment. The participants (coming from diverse programming backgrounds) were presented with pairs of code snippets (one using structural pattern matching while the other using traditional conditional branching), each addressing the same task. They had to choose which code they preferred using a 4- point Likert scale based on three criteria: readability, modifiability, and personal preference. In the vast majority of cases, developers preferred code with structural pattern matching, but there were certain contexts in which a significant proportion of developers preferred the original code. Norbert Vándor, Gabor Antal, Péter Hegedüs, Rudolf Ferenc |
SANER | 2 |
| 2022 | Don't DIY: Automatically transform legacy Python code to support structural pattern matchingabstractAs data becomes more and more complex as technology evolves, the need to support more complex data types in programming languages has grown. However, without proper storage and manipulation capabilities, handling such data can result in hard-to-read, difficult-to-maintain code. Therefore, programming languages continuously evolve to provide more and more ways to handle complex data. Python 3.10 introduced structural pattern matching, which serves this exact purpose: we can split complex data into relevant parts by examining its structure, and store them for later processing. Previously, we could only use the traditional conditional branching, which could have led to long chains of nested conditionals. Maintaining such code fragments can be cumbersome. In this paper, we present a complete framework to solve the aforementioned problem. Our software is capable of examining Python source code and transforming relevant conditionals into structural pattern matching. Moreover, it is able to handle nested conditionals and it is also easily extensible, thus the set of possible transformations can be easily increased. Balázs Rózsa, Gabor Antal, Rudolf Ferenc |
SCAM | 2 |
| 2021 | On the Rise and Fall of Simple Stupid Bugs: a Life-Cycle Analysis of SStuBsabstractBug detection and prevention is one of the most important goals of software quality assurance. Nowadays, many of the major problems faced by developers can be detected or even fixed fully or partially with automatic tools. However, recent works explored that there exists a substantial amount of simple yet very annoying errors in code-bases, which are easy to fix, but hard to detect as they do not hinder the functionality of the given product in a major way. Programmers introduce such errors accidentally, mostly due to inattention.Using the ManySStuBs4J dataset, which contains many simple, stupid bugs, found in GitHub repositories written in the Java programming language, we investigated the history of such bugs. We were interested in properties such as: How long do such bugs stay unnoticed in code-bases? Whether they are typically fixed by the same developer who introduced them? Are they introduced with the addition of new code or caused more by careless modification of existing code? We found that most of such stupid bugs lurk in the code for a long time before they get removed. We noticed that the developer who made the mistake seems to find a solution faster, however less then half of SStuBs are fixed by the same person. We also examined PMD's performance when to came to flagging lines containing SStuBs, and found that similarly to SpotBugs, it is insufficient when it comes to finding these types of errors. Examining the life-cycle of such bugs allows us to better understand their nature and adjust our development processes and quality assurance methods to better support avoiding them. Balázs Mosolygó, Norbert Vándor, Gabor Antal, Péter Hegedüs |
MSR | 3 |
| 2020 | A Data-Mining Based Study of Security Vulnerability Types and Their Mitigation in Different Languages
Gabor Antal, Balázs Mosolygó, Norbert Vándor, Péter Hegedüs |
ICCSA (4) | 1 |
| 2020 | Exploring the Security Awareness of the Python and JavaScript Open Source CommunitiesabstractSoftware security is undoubtedly a major concern in today's software engineering. Although the level of awareness of security issues is often high, practical experiences show that neither preventive actions nor reactions to possible issues are always addressed properly in reality. By analyzing large quantities of commits in the open-source communities, we can categorize the vulnerabilities mitigated by the developers and study their distribution, resolution time, etc. to learn and improve security management processes and practices. Gabor Antal, Márton Keleti, Péter Hegedüs |
MSR | 1 |
| 2018 | A Hands-on OpenStack Code Refactoring Experience Report
Gabor Antal, Alex Szarka, Péter Hegedüs |
ICCSA (5) | 1 |
| 2018 | [Research Paper] Static JavaScript Call Graphs: A Comparative StudyabstractThe popularity and wide adoption of JavaScript both at the client and server side makes its code analysis more important than ever before. Most of the algorithms for vulnerability analysis, coding issue detection, or type inference rely on the call graph representation of the underlying program. Despite some obvious advantages of dynamic analysis, static algorithms should also be considered for call graph construction as they do not require extensive test beds for programs and their costly execution and tracing. In this paper, we systematically compare five widely adopted static algorithms - implemented by the npm call graph, IBM WALA, Google Closure Compiler, Approximate Call Graph, and Type Analyzer for JavaScript tools - for building JavaScript call graphs on 26 WebKit SunSpider benchmark programs and 6 real-world Node.js modules. We provide a performance analysis as well as a quantitative and qualitative evaluation of the results. We found that there was a relatively large intersection of the found call edges among the algorithms, which proved to be 100% precise. However, most of the tools found edges that were missed by all others. ACG had the highest precision followed immediately by TAJS, but ACG found significantly more call edges. As for the combination of tools, ACG and TAJS together covered 99% of the found true edges by all algorithms, while maintaining a precision as high as 98%. Only two of the tools were able to analyze up-to-date multi-file Node.js modules due to incomplete language features support. They agreed on almost 60% of the call edges, but each of them found valid edges that the other missed. Gabor Antal, Péter Hegedüs, Zoltán Tóth, Rudolf Ferenc, Tibor Gyimóthy |
SCAM | 1 |
| 2017 | Empirical study on refactoring large-scale industrial systems and its effects on maintainability
Gábor Szoke, Gabor Antal, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy |
J. Syst. Softw. | 2 |
| 2016 | Transforming C++11 Code to C++03 to Support Legacy Compilation EnvironmentsabstractNewer technologies - programming languages, environments, libraries - change very rapidly. However, various internal and external constraints often prevent projects from quickly adopting to these changes. Customers may require specific platform compatibility from a software vendor, for example. In this work, we deal with such an issue in the context of the C++ programming language. Our industrial partner is required to use SDKs that support only older C++ language editions. They, however, would like to allow their developers to use the newest language constructs in their code. To address this problem, we created a source code transformation framework to automatically backport source code written according to the C++11 standard to its functionally equivalent C++03 variant. With our framework developers are free to exploit the latest language features, while production code is still built by using a restricted set of available language constructs. This paper reports on the technical details of the transformation engine, and our experiences in applying it on two large industrial code bases and four open-source systems. Our solution is freely available and open-source. Gabor Antal, David Havas, István Siket, Árpád Beszédes, Rudolf Ferenc, József Mihalicza |
SCAM | 1 |
| 2015 | Identifying wasted effort in the field via developer interaction dataabstractDuring software projects, several parts of the source code are usually re-written due to imperfect solutions before the code is released. This wasted effort is of central interest to the project management to assure on-time delivery. Although the amount of thrown-away code can be measured from version control systems, stakeholders are more interested in productivity dynamics that reflect the constant change in a software project. In this paper we present a field study of measuring the productivity of a medium-sized J2EE project. We propose a productivity analysis method where productivity is expressed through dynamic profiles - the so-called Micro-Productivity Profiles (MPPs). They can be used to characterize various constituents of software projects such as components, phases and teams. We collected detailed traces of developers' actions using an Eclipse IDE plug-in for seven months of software development throughout two milestones. We present and evaluate profiles of two important axes of the development process: by milestone and by application layers. MPPs can be an aid to take project control actions and help in planning future projects. Based on the experiments, project stakeholders identified several points to improve the development process. It is also acknowledged, that profiles show additional information compared to a naive diff-based approach. Gergö Balogh, Gabor Antal, Árpád Beszédes, László Vidács, Tibor Gyimóthy, Ádám Zoltán Végh |
ICSME | 2 |
| 2014 | Bulk Fixing Coding Issues and Its Effects on Software Quality: Is It Worth Refactoring?abstractThe quality of a software system is mostly defined by its source code. Software evolves continuously, it gets modified, enhanced, and new requirements always arise. If we do not spend time periodically on improving our source code, it becomes messy and its quality will decrease inevitably. Literature tells us that we can improve the quality of our software product by regularly refactoring it. But does refactoring really increase software quality? Can it happen that a refactoring decreases the quality? Is it possible to recognize the change in quality caused by a single refactoring operation? In our paper, we seek answers to these questions in a case study of refactoring large-scale proprietary software systems. We analyzed the source code of 5 systems, and measured the quality of several revisions for a period of time. We analyzed 2 million lines of code and identified nearly 200 refactoring commits which fixed over 500 coding issues. We found that one single refactoring only makes a small change (sometimes even decreases quality), but when we do them in blocks, we can significantly increase quality, which can result not only in the local, but also in the global improvement of the code. Gábor Szoke, Gabor Antal, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy |
SCAM | 2 |