VLDB 2026 Research / reviewers in the wild / expert
Péter Hegedüs
dblp:33/1061
· DBLP profile ↗
28ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-4592-6504ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary StudyabstractRecent advancements in large language models (LLMs) have shown promise for automated vulnerability detection and repair in software systems. This paper investigates the performance of GPT-4o in repairing Java vulnerabilities from a widely used dataset (Vul4J), exploring how different contextual information affects automated vulnerability repair (AVR) capabilities. We compare the latest GPT-4o’s performance against previous results with GPT-4 using identical prompts. We evaluated nine additional prompts crafted by us that contain various contextual information such as CWE or CVE information, and manually extracted code contexts. Each prompt was executed three times on 42 vulnerabilities, and the resulting fix candidates were validated using Vul4J’s automated testing framework. Our results show that GPT-4o performed 11.9% worse on average than GPT-4 with the same prompt, but was able to fix 10.5% more distinct vulnerabilities in the three runs together. CVE information significantly improved repair rates, while the length of the task description had minimal impact. Combining CVE guidance with manually extracted code context resulted in the best performance. Using our Top-3 prompts together, GPT-4o repaired 26 (62%) vulnerabilities at least once, outperforming both the original baseline (40%) and its reproduction (45%), suggesting that ensemble prompt strategies could improve vulnerability repair in zero-shot settings. Gabor Antal, Bence Bogenfürst, Rudolf Ferenc, Péter Hegedüs |
EASE | 4 |
| 2025 | Leveraging GPT-4 for Vulnerability-Witnessing Unit Test GenerationabstractIn the life-cycle of software development, testing plays a crucial role in quality assurance. Proper testing not only increases code coverage and prevents regressions but it can also ensure that any potential vulnerabilities in the software are identified and effectively fixed. However, creating such tests is a complex, resource-consuming manual process. To help developers and security experts, this paper explores the automatic unit test generation capability of one of the most widely used large language models, GPT-4, from the perspective of vulnerabilities. We examine a subset of the VUL4J dataset containing real vulnerabilities and their corresponding fixes to determine whether GPT-4 can generate syntactically and/or semantically correct unit tests based on the code before and after the fixes as evidence of vulnerability mitigation. We focus on the impact of code contexts, the effectiveness of GPT-4’s self-correction ability, and the subjective usability of the generated test cases. Our results indicate that GPT-4 can generate syntactically correct test cases 66.5% of the time without domain-specific pre-training. Although the semantic correctness of the fixes could be automatically validated in only 7. 5% of the cases, our subjective evaluation shows that GPT-4 generally produces test templates that can be further developed into fully functional vulnerability-witnessing tests with relatively minimal manual effort. Gabor Antal, Dénes Bán, Martin Isztin, Rudolf Ferenc, Péter Hegedüs |
EASE | 5 |
| 2024 | Reality Check: Assessing GPT-4 in Fixing Real-World Software VulnerabilitiesabstractDiscovering and mitigating software vulnerabilities is a challenging task. These vulnerabilities are often caused by simple, otherwise (and in other contexts) harmless code snippets (e.g., unchecked path traversal). Large Language Models (LLMs) promise to revolutionize not just human-machine interactions but various software engineering tasks as well, including the automatic repair of vulnerabilities. However, currently, it is hard to assess the performance, robustness, and reliability of these models as most of their evaluation has been done on small, synthetic examples. In our work, we systematically evaluate the automatic vulnerability fixing capabilities of GPT-4, a popular LLM, using a database of real-world Java vulnerabilities, Vul4J. We expect the model to provide fixes for vulnerable methods, which we evaluate manually and based on unit test results included in the Vul4J database. GPT-4 provided perfect fixes consistently for at least 12 out of the total 46 examined vulnerabilities, which could be applied as is. In an additional 5 cases, the provided textual instructions would help to fix the vulnerabilities in a practical scenario (despite the provided code being incorrect). Our findings, similar to others, also show that prompting has a significant effect. Zoltán Ságodi, Gabor Antal, Bence Bogenfürst, Martin Isztin, Péter Hegedüs, Rudolf Ferenc |
EASE | 5 |
| 2024 | On the Usefulness of Python Structural Pattern Matching: An Empirical StudyabstractAs the important role of software in our modern world becomes more and more evident, the need for more complex data structures is increasing. Structural pattern matching has become an elegant technique for simplifying complex conditionallogic in the source code in modern programming languages. It provides an easy and concise way to destructure complex data and enables making decisions based on its structure, ultimately improving code readability. In the context of Python, a language renowned for its simplicity, structural pattern matching was a missing feature until October 2021. Finally, the feature is shipped in Python 3.10, allowing developers to exploit the potential of structural pattern matching. Instead of traditional conditional branching and endless type checking, structural pattern matching offers a more straightforward way to handle data. In this paper, we investigate the usefulness of this relatively new Python feature by involving 65 participants in a code review experiment. The participants (coming from diverse programming backgrounds) were presented with pairs of code snippets (one using structural pattern matching while the other using traditional conditional branching), each addressing the same task. They had to choose which code they preferred using a 4- point Likert scale based on three criteria: readability, modifiability, and personal preference. In the vast majority of cases, developers preferred code with structural pattern matching, but there were certain contexts in which a significant proportion of developers preferred the original code. Norbert Vándor, Gabor Antal, Péter Hegedüs, Rudolf Ferenc |
SANER | 3 |
| 2022 | A Vulnerability Introducing Commit Dataset for Java: An Improved SZZ based ApproachabstractIn the domain of vulnerability detection from the source code by applying static analysis, the number and quality of available datasets for creating and testing security analysis methods is quite low.To be precise, there are already several public datasets containing vulnerability fixing commits; however, vulnerability introducing commit datasets are scarce, which would be essential for creating and validating just-in-time vulnerability detection approaches.In this paper, we propose an SZZ (an algorithm originally developed to find bug introducing commits) based method with a specific filtering mechanism to create vulnerability introducing commit datasets from vulnerability fixes.The filtering phase involves measuring a relevance score for each vulnerability introducing commit candidates based on commit similarities.We generated a novel Java vulnerability introducing dataset from the existing project-KB repository to demonstrate our algorithm's capabilities.We also showcase the generated database and the effectiveness of our filtering method through several hand-picked examples from the dataset. INTRODUCTIONMany software engineering-related tasks, such as quality assurance or testing, are now aided by machine learning, which relies heavily on the abundance of data.Most of these tasks are typically based on machine learning, therefore the availability of datasets is crucial to train reliably and to get a generally wellperforming model.Fortunately, when the goal is related to vulnerability fixes, there are already well established datasets that can be relied on.These datasets typically contain validated code changes (i.e.commits) that fix a particular vulnerability described in a Common Vulnerabilities and Exposures (CVE) (MITRE Corporation, v 21) entry, a publicly disclosed security vulnerability in a software system.One such dataset is published as part of the repository "project-KB" (project kaybee) (Ponta et al., 2019) maintained by SAP. Tamás Aladics, Péter Hegedüs, Rudolf Ferenc |
ICSOFT | 2 |
| 2022 | Is Refactoring Always a Good Egg? Exploring the Interconnection Between Bugs and RefactoringsabstractBug fixing and code refactoring are two distinct maintenance actions with different goals. While bug fixing is a corrective change that eliminates a defect from the program, refactoring targets improving the internal quality (i.e., maintainability) of a software system without changing its functionality. Best practices and common intuition suggest that these code actions should not be mixed in a single code change. Furthermore, as refactoring aims for improving quality without functional changes, we would expect that refactoring code changes will not be sources of bugs. Nonetheless, empirical studies show that none of the above hypotheses are necessarily true in practice. In this paper, we empirically investigate the interconnection between bug-related and refactoring code changes using the SmartSHARK dataset. Our goal is to explore how often bug fixes and refactorings co-occur in a single commit (tangled changes) and whether refactoring changes themselves might induce bugs into the system. We found that it is not uncommon to have tangled commits of bug fixes and refactorings; 21% of bug-fixing commits include at least one type of refactoring on average. What is even more shocking is that 54% of bug-inducing commits also contain code refactoring changes. For instance, 10% (652 occurrences) of the Change Variable Type refactorings in the dataset appear in bug-inducing commits that make up 7.9% of the total inducing commits. Amirreza Bagheri, Péter Hegedüs |
MSR | 2 |
| 2022 | An End-to-End Framework for Repairing Potentially Vulnerable Source CodeabstractNowadays, program development is getting easier and easier as the various IDE tools provide advice on what to write in the program. But it is not enough to implement a solution to a problem; it is also important that the non-functional properties, like the quality or security of the code, are appropriate in all aspects. One of the most widely used techniques to ensure quality is testing. If the tests fail, one can fix the code immediately. However, security issues are unexpected cases when implementing the program, which is why we do not write tests for them in advance. In many cases, security-relevant bugs can not only cause financial loss but also put human lives at risk, so detecting and fixing them is an important step for the reliability and quality of the program. The tool presented in this paper aims to generate automatic code repairs to potential vulnerabilities in the program. By integrating the recommended fixes, one can easily harden the security of their program early in the development process. A case study on six open-source Java subject systems showed that we were able to generate viable repair patches for 57 out of the 81 detected security issues (70%). For certain types (e.g., revealing private references of mutable objects), our tool reached close to perfect performance. Judit Jász, Péter Hegedüs, Ákos Milánkovich, Rudolf Ferenc |
SCAM | 2 |
| 2021 | Machine Learning Applied for Spectra Classification
Sandor Brockhauser, Péter Hegedüs |
ICCSA (9) | 3 |
| 2021 | Improving Vulnerability Prediction of JavaScript Functions using Process MetricsabstractDue to the growing number of cyber attacks against computer systems, we need to pay special attention to the security of our software systems. In order to maximize the effectiveness, excluding the human component from this process would be a huge breakthrough. The first step towards this is to automatically recognize the vulnerable parts in our code. Researchers put a lot of effort into creating machine learning models that could determine if a given piece of code, or to be more precise, a selected function, contains any vulnerabilities or not. We aim at improving the existing models, building on previous results in predicting vulnerabilities at the level of functions in JavaScript code using the well-known static source code metrics. In this work, we propose to include several so-called process metrics (e.g., code churn, number of developers modifying a file, or the age of the changed source code) into the set of features, and examine how they affect the performance of the function-level JavaScript vulnerability prediction models. We can confirm that process metrics significantly improve the prediction power of such models. On average, we observed a 8.4% improvement in terms of F-measure (from 0.764 to 0.848), 3.5% improvement in terms of precision (from 0.953 to 0.988) and a 6.3% improvement in terms of recall (from 0.697 to 0.760). Tamás Viszkok, Péter Hegedüs, Rudolf Ferenc |
ICSOFT | 2 |
| 2021 | On the Rise and Fall of Simple Stupid Bugs: a Life-Cycle Analysis of SStuBsabstractBug detection and prevention is one of the most important goals of software quality assurance. Nowadays, many of the major problems faced by developers can be detected or even fixed fully or partially with automatic tools. However, recent works explored that there exists a substantial amount of simple yet very annoying errors in code-bases, which are easy to fix, but hard to detect as they do not hinder the functionality of the given product in a major way. Programmers introduce such errors accidentally, mostly due to inattention.Using the ManySStuBs4J dataset, which contains many simple, stupid bugs, found in GitHub repositories written in the Java programming language, we investigated the history of such bugs. We were interested in properties such as: How long do such bugs stay unnoticed in code-bases? Whether they are typically fixed by the same developer who introduced them? Are they introduced with the addition of new code or caused more by careless modification of existing code? We found that most of such stupid bugs lurk in the code for a long time before they get removed. We noticed that the developer who made the mistake seems to find a solution faster, however less then half of SStuBs are fixed by the same person. We also examined PMD's performance when to came to flagging lines containing SStuBs, and found that similarly to SpotBugs, it is insufficient when it comes to finding these types of errors. Examining the life-cycle of such bugs allows us to better understand their nature and adjust our development processes and quality assurance methods to better support avoiding them. Balázs Mosolygó, Norbert Vándor, Gabor Antal, Péter Hegedüs |
MSR | 4 |
| 2020 | A Data-Mining Based Study of Security Vulnerability Types and Their Mitigation in Different Languages
Gabor Antal, Balázs Mosolygó, Norbert Vándor, Péter Hegedüs |
ICCSA (4) | 4 |
| 2020 | Inspecting JavaScript Vulnerability Mitigation Patches with Automated Fix Generation in Mind
Péter Hegedüs |
ICCSA (4) | 1 |
| 2020 | Exploring the Security Awareness of the Python and JavaScript Open Source CommunitiesabstractSoftware security is undoubtedly a major concern in today's software engineering. Although the level of awareness of security issues is often high, practical experiences show that neither preventive actions nor reactions to possible issues are always addressed properly in reality. By analyzing large quantities of commits in the open-source communities, we can categorize the vulnerabilities mitigated by the developers and study their distribution, resolution time, etc. to learn and improve security management processes and practices. Gabor Antal, Márton Keleti, Péter Hegedüs |
MSR | 3 |
| 2018 | A Hands-on OpenStack Code Refactoring Experience Report
Gabor Antal, Alex Szarka, Péter Hegedüs |
ICCSA (5) | 3 |
| 2018 | Developer Focus: Lack of Impact on Maintainability
Csaba Faragó, Péter Hegedüs |
ICCSA (5) | 2 |
| 2018 | [Research Paper] Static JavaScript Call Graphs: A Comparative StudyabstractThe popularity and wide adoption of JavaScript both at the client and server side makes its code analysis more important than ever before. Most of the algorithms for vulnerability analysis, coding issue detection, or type inference rely on the call graph representation of the underlying program. Despite some obvious advantages of dynamic analysis, static algorithms should also be considered for call graph construction as they do not require extensive test beds for programs and their costly execution and tracing. In this paper, we systematically compare five widely adopted static algorithms - implemented by the npm call graph, IBM WALA, Google Closure Compiler, Approximate Call Graph, and Type Analyzer for JavaScript tools - for building JavaScript call graphs on 26 WebKit SunSpider benchmark programs and 6 real-world Node.js modules. We provide a performance analysis as well as a quantitative and qualitative evaluation of the results. We found that there was a relatively large intersection of the found call edges among the algorithms, which proved to be 100% precise. However, most of the tools found edges that were missed by all others. ACG had the highest precision followed immediately by TAJS, but ACG found significantly more call edges. As for the combination of tools, ACG and TAJS together covered 99% of the found true edges by all algorithms, while maintaining a precision as high as 98%. Only two of the tools were able to analyze up-to-date multi-file Node.js modules due to incomplete language features support. They agreed on almost 60% of the call edges, but each of them found valid edges that the other missed. Gabor Antal, Péter Hegedüs, Zoltán Tóth, Rudolf Ferenc, Tibor Gyimóthy |
SCAM | 2 |
| 2018 | Empirical evaluation of software maintainability based on a manually validated refactoring dataset
Péter Hegedüs, István Kádár, Rudolf Ferenc, Tibor Gyimóthy |
Inf. Softw. Technol. | 1 |
| 2016 | Assessment of the Code Refactoring Dataset Regarding the Maintainability of Methods
István Kádár, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy |
ICCSA (4) | 2 |
| 2016 | A Code Refactoring Dataset and Its Assessment Regarding Software MaintainabilityabstractIt is very common in various fields that there is a gap between theoretical results and their practical applications. This is true for code refactoring as well, which has a solid theoretical background while being used in development practice at the same time. However, more and more studies suggest that developers perform code refactoring entirely differently than the theory would suggest. Our paper encourages the further investigation of code refactorings in practice by providing an excessive open dataset of source code metrics and applied refactorings through several releases of 7 open-source systems. As a first step of processing this dataset, we examined the quality attributes of the refactored source code classes and the values of source code metrics improved by those refactorings. Our early results show that lower maintainability indeed triggers more code refactorings in practice and these refactorings significantly decrease complexity, code lines, coupling and clone metrics. However, we observed a decrease in comment related metrics in the refactored code. István Kádár, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy |
SANER | 2 |
| 2015 | Code Ownership: Impact on Maintainability
Csaba Faragó, Péter Hegedüs, Rudolf Ferenc |
ICCSA (5) | 2 |
| 2015 | Adding Constraint Building Mechanisms to a Symbolic Execution Engine Developed for Detecting Runtime Errors
István Kádár, Péter Hegedüs, Rudolf Ferenc |
ICCSA (5) | 2 |
| 2015 | Advances in software product quality measurement and its applications in software evolutionabstractThe main results presented in this work, a synopsis of the connected PhD dissertation, are related to software product quality modeling and measurement as well as to the application of the newly proposed methods, tools and techniques in software evolution. All the novel theoretical results and models were thoroughly validated via empirical case studies and successfully applied in practice. The thesis result statements can be grouped into three major points: (i) system-level software quality models; (ii) source code element-level software quality models; (iii) applications of the proposed quality models. Some of the methods and tools presented in the thesis have been utilized in Hungarian and international R&D projects as well as by the industrial partners of the Software Engineering Department of the University of Szeged. Péter Hegedüs |
ICSME | 1 |
| 2015 | Do automatic refactorings improve maintainability? An industrial case studyabstractRefactoring is often treated as the main remedy against the unavoidable code erosion happening during software evolution. Studies show that refactoring is indeed an elemental part of the developers' arsenal. However, empirical studies about the impact of refactorings on software maintainability still did not reach a consensus. Moreover, most of these empirical investigations are carried out on open-source projects where distinguishing refactoring operations from other development activities is a challenge in itself. We had a chance to work together with several software development companies in a project where they got extra budget to improve their source code by performing refactoring operations. Taking advantage of this controlled environment, we collected a large amount of data during a refactoring phase where the developers used a (semi)automatic refactoring tool. By measuring the maintainability of the involved subject systems before and after the refactorings, we got valuable insights into the effect of these refactorings on large-scale industrial projects. All but one company, who applied a special refactoring strategy, achieved a maintainability improvement at the end of the refactoring phase, but even that one company suffered from the negative impact of only one type of refactoring. Gábor Szoke, Csaba Nagy 0001, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy |
ICSME | 3 |
| 2015 | Cumulative code churn: Impact on maintainabilityabstractIt is a well-known phenomena that the source code of software systems erodes during development, which results in higher maintenance costs in the long term. But can we somehow narrow down where exactly this erosion happens? Is it possible to infer the future erosion based on past code changes? Do modifications performed on frequently changing code have worse effect on software maintainability than those affecting less frequently modified code? In this study we investigated these questions and the results indicate that code churn indeed increases the pace of code erosion. We calculated cumulative code churn values and maintainability changes for every version control commit operation of three open-source and one proprietary software system. With the help of Wilcoxon rank test we compared the cumulative code churn values of the files in commits resulting maintainability increase with those of decreasing the maintainability. In the case of three systems the test showed very strong significance and in one case it resulted in strong significance (p-values 0.00235, 0.00436, 0.00018 and 0.03616). These results support our preliminary assumption that modifying high-churn code is more likely to decrease the overall maintainability of a software system, which can be thought of as the generalization of the already known phenomena that code churn results in higher number of defects. Csaba Faragó, Péter Hegedüs, Rudolf Ferenc |
SCAM | 2 |
| 2014 | The Impact of Version Control Operations on the Quality Change of the Source Code
Csaba Faragó, Péter Hegedüs, Rudolf Ferenc |
ICCSA (5) | 2 |
| 2013 | Revealing the Effect of Coding Practices on Software MaintainabilityabstractDue to its very obvious and direct connection with the costs of altering the behavior of a software, maintainability is probably the most attractive, observed and evaluated quality characteristic of the software products. There are many coding practices and techniques that may influence the maintainability of a system (e.g. design patterns, coding rules, anti-patterns, refactoring techniques). However, the empirical evidences of the connection between coding practices and maintainability are vague due to the following reasons: i) finding instances of coding primitives like design patterns, anti-patterns, etc. precisely with reverse engineering tools is not easy, ii) the lack of mature practical quality models for objective calculation of maintainability and handling its ambiguity, iii) few empirical studies directly evaluating the connection of coding techniques and software maintainability. The presented work focuses on solving these major problems by creating a benchmark for evaluating the performance of different reverse engineering tools and introducing a novel probabilistic approach for measuring software maintainability. By performing case studies based on new analysis methods we evince that there is a significant correlation between the design pattern density and the maintainability of a system, e.g. 0.89 Pearson correlation for JHotDraw. Moreover, preliminary studies show that applying refactoring has indeed a traceable positive impact on software maintainability as anticipated. Péter Hegedüs |
ICSM | 1 |
| 2012 | A cost model based on software maintainabilityabstractIn this paper we present a maintainability based model for estimating the costs of developing source code in its evolution phase. Our model adopts the concept of entropy in thermodynamics, which is used to measure the disorder of a system. In our model, we use maintainability for measuring disorder (i.e. entropy) of the source code of a software system. We evaluated our model on three proprietary and two open source real world software systems implemented in Java, and found that the maintainability of these evolving software is decreasing over time. Furthermore, maintainability and development costs are in exponential relationship with each other. We also found that our model is able to predict future development costs with high accuracy in these systems. Tibor Bakota, Péter Hegedüs, Gergely Ladányi, Peter Kortvelyesi, Rudolf Ferenc, Tibor Gyimóthy |
ICSM | 2 |
| 2011 | A probabilistic software quality modelabstractIn order to take the right decisions in estimating the costs and risks of a software change, it is crucial for the developers and managers to be aware of the quality attributes of their software. Maintainability is an important characteristic defined in the ISO/IEC 9126 standard, owing to its direct impact on development costs. Although the standard provides definitions for the quality characteristics, it does not define how they should be computed. Not being tangible notions, these characteristics are hardly expected to be representable by a single number. Existing quality models do not deal with ambiguity coming from subjective interpretations of characteristics, which depend on experience, knowledge, and even intuition of experts. This research aims at providing a probabilistic approach for computing high-level quality characteristics, which integrate expert knowledge, and deal with ambiguity at the same time. The presented method copes with “goodness” functions, which are continuous generalizations of threshold based approaches, i.e. instead of giving a number for the measure of goodness, it provides a continuous function. Two different systems were evaluated using this approach, and the results were compared to the opinions of experts involved in the development. The results show that the quality model values change in accordance with the maintenance activities, and they are in a good correlation with the experts' expectations. Tibor Bakota, Péter Hegedüs, Peter Kortvelyesi, Rudolf Ferenc, Tibor Gyimóthy |
ICSM | 2 |