VLDB 2026 Research / reviewers in the wild / expert
Ferenc Horváth
dblp:150/7797
· DBLP profile ↗
14ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-8442-7970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Context Switch Sensitive Fault LocalizationabstractSpectrum-Based Fault Localization (SBFL) is a popular technique to assist developers in pinpointing faulty elements within their code based on test outcomes and code coverage. In this paper, we examine the impact of context switching, i.e., when developers must frequently shift their attention between different code parts (such as methods and classes) while going down the SBFL ranked list to find the faulty statement. The basis of our study is the observation that it requires less effort to investigate statements that are next to each other rather than those in different methods and classes. In particular, we analyse the number of visited methods and classes, as well as the frequency of switches between them during the fault localization process. We found that, in programs from the Defects4J benchmark, developers need to explore 40 methods and 12 classes on average, before finding the faulty statement, leading to 53 method- and 40 class switches, respectively. Ferenc Horváth, Roland Aszmann, Péter Attila Soha, Árpád Beszédes, Tibor Gyimóthy |
EASE | 1 |
| 2024 | On the Stability and Applicability of Deep Learning in Fault LocalizationabstractNumerous Deep Learning (DL)-based fault localization (FL) methods are developed with the aim of leveraging the code coverage matrix and failure vector to identify the connection between program elements and defects. The imbalanced data on which these approaches train their models poses a substantial challenge to the effectiveness of fault localization techniques. This study explores the stability of fault localization models in deep learning, specifically, their performance when trained repeatedly using the same input but varying random initializations. Using the Defect4J benchmark, we trained deep learning models (MLP, CNN, and RNN) independently and found that 86 cases resulted in (partly) consistent rankings among all five models and versions, while 621 exhibited varying outcomes, meaning that 90 % of the produced ranks were different in subsequent trainings. The models showed significant variability in ranking results, with maximum ranks sometimes five times that of the minimum. We also adapted the churn metric from DL research to evaluate models, confirming their instability. To improve stability, meta-parameter optimization, model simplification and resampling has been applied. Although some of these techniques proved effective, even with the improvements, the models remained insufficiently stable to produce reliable results. Viktor Csuvik, Roland Aszmann, Árpád Beszédes, Ferenc Horváth, Tibor Gyimóthy |
SANER | 4 |
| 2023 | A Case Against Coverage-Based Program SpectraabstractSpectrum-Based Fault Localization (SBFL) is a semi-automated debugging technique that gained popularity in the last decades due to its intuitive approach and relatively simple implementability. Despite this, the performance of practical SBFL techniques in terms of fault localization capability does not reach the threshold that would enable their acceptance by professional programmers. Almost all modern SBFL approaches are based on the code coverage-based spectrum, and on the assumption that a code element covered by failing tests should be treated as suspicious. However, it is easy to see that this is an over-approximation because many code elements may be executed that do not contribute to the test output, hence serving as noise in the process. A possible solution is to use backward dynamic program slices as program spectra computed from the output statement as the criterion, instead of the coverage. There are very few theoretical and practical results about this approach, so in this work we revisit the method and show how much more inferior coverage-based spectra are compared to slice-based spectra, both on theoretical and practical levels. We argue that code coverage-based SBFL is currently in a research pit due to this inherent approximation, and research on slice-based spectra should once more attain a much higher focus. Péter Attila Soha, Tamás Gergely, Ferenc Horváth, Béla Vancsics, Árpád Beszédes |
ICST | 3 |
| 2022 | Using contextual knowledge in interactive fault localizationabstractAbstract Tool support for automated fault localization in program debugging is limited because state-of-the-art algorithms often fail to provide efficient help to the user. They usually offer a ranked list of suspicious code elements, but the fault is not guaranteed to be found among the highest ranks. In Spectrum-Based Fault Localization (SBFL) – which uses code coverage information of test cases and their execution outcomes to calculate the ranks –, the developer has to investigate several locations before finding the faulty code element. Yet, all the knowledge she a priori has or acquires during this process is not reused by the SBFL tool. There are existing approaches in which the developer interacts with the SBFL algorithm by giving feedback on the elements of the prioritized list. We propose a new approach called iFL which extends interactive approaches by exploiting contextual knowledge of the user about the next item in the ranked list (e. g., a statement), with which larger code entities (e. g., a whole function) can be repositioned in their suspiciousness. We implemented a closely related algorithm proposed by Gong et al., called Talk. First, we evaluated iFL using simulated users, and compared the results to SBFL and Talk. Next, we introduced two types of imperfections in the simulation: user’s knowledge and confidence levels. On SIR and Defects4J, results showed notable improvements in fault localization efficiency, even with strong user imperfections. We then empirically evaluated the effectiveness of the approach with real users in two sets of experiments: a quantitative evaluation of the successfulness of using iFL, and a qualitative evaluation of practical uses of the approach with experienced developers in think-aloud sessions. Ferenc Horváth, Árpád Beszédes, Béla Vancsics, Gergö Balogh, László Vidács, Tibor Gyimóthy |
Empir. Softw. Eng. | 1 |
| 2022 | Fault localization using function call frequenciesabstractIn traditional Spectrum-Based Fault Localization (SBFL), hit-based spectrum is used to estimate a program element’s suspiciousness to contain a fault, i.e., only the binary information is used if the code element was executed by the test case or not. Count-based spectra can potentially improve the localization effectiveness due to the number of executions also being available. In this work, we use function-level granularity and define count-based spectra which use function call frequencies. We investigate the naïve approach, which simply counts the function call instances. We also define a novel method which is based on counting the different function call contexts, i.e., the frequency of the investigated function occurring in unique call stack instances during test execution. The basic intuition is that if a function is called in many different contexts during a failing test case, it will be more probable to be accountable for the fault. We empirically evaluated the fault localization capability of different variations of the approach and compared them to 9 traditional SBFL techniques using the Defects4J benchmark. We show that: (i) naïve counts result in worse rank positions than the hit-based approach, but (ii) unique counts produce better rank positions with some of the algorithm variants. Béla Vancsics, Ferenc Horváth, Attila Szatmári, Árpád Beszédes |
J. Syst. Softw. | 2 |
| 2021 | Call Frequency-Based Fault LocalizationabstractSpectrum-Based Fault Localization (SBFL), in its basic form, uses only local information about a program element’s (such as a method’s) coverage to predict its faultiness, and rarely is any additional (contextual) information leveraged about the element itself, nor the test cases. As such an additional context, in the presented approach, we rely on the frequency of the investigated method occurring in call stack instances during the course of executing the failing test cases. The basic intuition is that if a method is called in many different contexts during a failing test case, it will be more probable to be accountable for the fault compared to other methods. We empirically evaluated the fault localization capability of the approach compared to five traditional SBFL techniques using the bug benchmark Defects4J. We found that the new algorithms (i) find the location of bugs at higher rank positions more often, (ii) can achieve 38%–52% rank position improvement compared to the baseline algorithms with statistical significance, and (iii) place more items at the top-10 positions of the suspiciousness ranking. Béla Vancsics, Ferenc Horváth, Attila Szatmári, Árpád Beszédes |
SANER | 2 |
| 2020 | Experiments with Interactive Fault Localization Using Simulated and Real UsersabstractFault localization is considered a difficult and time consuming activity. However, tool support for automated fault localization is still limited because state-of-the-art algorithms often fail to provide efficient help to the user. They usually offer a ranked list of suspicious code elements, but the fault is not guaranteed to be found among the highest ranks. In Spectrum-Based Fault Localization (SBFL) - which uses code coverage information of test cases and their execution outcomes to calculate the ranks -, the developer has to investigate several locations before finding the faulty code element. Yet, all the knowledge she a priori has or acquires during this process is not reused by the SBFL tool. We propose an approach in which the developer interacts with the SBFL algorithm by giving feedback on the elements of the prioritized list. We exploit contextual knowledge of the user about the next item in the ranked list (e. g., a statement), with which larger code entities (e. g., a whole function) can be repositioned in their suspiciousness. First, we evaluated the approach using simulated users incorporating two types of imperfections, their knowledge and confidence levels. On SIR and Defects4J, results showed notable improvements in fault localization efficiency, even with strong user imperfections. We then empirically evaluated the effectiveness of the approach with real users, which also showed promising results. Ferenc Horváth, Árpád Beszédes, Béla Vancsics, Gergö Balogh, László Vidács, Tibor Gyimóthy |
ICSME | 1 |
| 2020 | Leveraging Contextual Information from Function Call Chains to Improve Fault LocalizationabstractIn Spectrum Based Fault Localization, program elements such as statements or functions are ranked according to a suspiciousness score which can guide the programmer in finding the fault more efficiently. However, such a ranking does not include any additional information about the suspicious code elements. In this work, we propose to complement function-level spectrum based fault localization with function call chains - i.e., snapshots of the call stack occurring during execution - on which the fault localization is first performed, and then narrowed down to functions. Our experiments using defects from Defects4J show that (i) 69% of the defective functions can be found in call chains with highest scores, (ii) in 4 out of 6 cases the proposed approach can improve Ochiai ranking of 1 to 9 positions on average, with a relative improvement of 19–48%, and (iii) the improvement is substantial (66–98%) when Ochiai produces bad rankings for the faulty functions. Árpád Beszédes, Ferenc Horváth, Massimiliano Di Penta, Tibor Gyimóthy |
SANER | 2 |
| 2019 | Poster: Aiding Java Developers with Interactive Fault Localization in Eclipse IDEabstractSpectrum-Based ones are a popular class of Fault Localization (FL) methods among researchers due to their relative simplicity. However, recent studies highlighted some barriers to the wider adoption of the technique in practical settings. One possibility to increase the practical usefulness of related tools is to involve interactivity between the user and the core FL algorithm. In this setting, the developer interacts with the fault localization algorithm by giving feedback on the elements proposed by the algorithm. This way, the proposed elements can be influenced in the hope to reach the faulty element earlier (we call the proposed approach Interactive Fault Localization, or iFL). With this work, we present our recent achievements in this topic. In particular, we overview the basic approach, our preliminary experimentation with user simulation, and the supporting tool for the actual usage of the method, iFL for Eclipse. Our aim is to provide a basis for the investigation of the feasibility and effectiveness of the technique, before moving on to more comprehensive experiments with actual human subjects. We invite researchers for further discussion on the topic, and for that, the method and tool will be made accessible. Gergö Balogh, Ferenc Horváth, Árpád Beszédes |
ICST | 2 |
| 2019 | Feature analysis using information retrieval, community detection and structural analysis methods in product line adoptionabstractIn industrial practice the clone-and-own strategy is often applied when in the pressure of high demand of customized features. The adoption of software product line (SPL) architecture is a large one time investment that affects both technical and organizational issues. The analysis of the feature structure is a crucial point in the SPL adoption process involving domain experts working at a higher level of abstraction and developers working directly on the program code. We propose automatic methods to extract feature-to-program links starting from very high level set of features provided by domain experts. For this purpose we combine call graph information with textual similarity between code and high level features. In addition, in depth understanding of the feature structure is supported by finding communities between programs and relating them to features. As features are originated from domain experts, community analysis reveals discrepancies between expert view and internal code structure. We found that communities correspond well to the high level features, with usually more than half of feature code located in specialized communities. We report experiments at two levels of features and more than 2000 Magic 4GL programs in an industrial SPL adoption project. András Kicsi, Viktor Csuvik, László Vidács, Ferenc Horváth, Árpád Beszédes, Tibor Gyimóthy, Ferenc Kocsis |
J. Syst. Softw. | 4 |
| 2019 | Differences between a static and a dynamic test-to-code traceability recovery methodabstractRecovering test-to-code traceability links may be required in virtually every phase of development. This task might seem simple for unit tests thanks to two fundamental unit testing guidelines: isolation (unit tests should exercise only a single unit) and separation (they should be placed next to this unit). However, practice shows that recovery may be challenging because the guidelines typically cannot be fully followed. Furthermore, previous works have already demonstrated that fully automatic test-to-code traceability recovery for unit tests is virtually impossible in a general case. In this work, we propose a semi-automatic method for this task, which is based on computing traceability links using static and dynamic approaches, comparing their results and presenting the discrepancies to the user, who will determine the final traceability links based on the differences and contextual information. We define a set of discrepancy patterns, which can help the user in this task. Additional outcomes of analyzing the discrepancies are structural unit testing issues and related refactoring suggestions. For the static test-to-code traceability, we rely on the physical code structure, while for the dynamic, we use code coverage information. In both cases, we compute combined test and code clusters which represent sets of mutually traceable elements. We also present an empirical study of the method involving 8 non-trivial open source Java systems. Tamás Gergely, Gergö Balogh, Ferenc Horváth, Béla Vancsics, Árpád Beszédes, Tibor Gyimóthy |
Softw. Qual. J. | 3 |
| 2019 | Code coverage differences of Java bytecode and source code instrumentation tools
Ferenc Horváth, Tamás Gergely, Árpád Beszédes, Dávid Tengeri, Gergö Balogh, Tibor Gyimóthy |
Softw. Qual. J. | 1 |
| 2018 | Supporting Product Line Adoption by Combining Syntactic and Textual Feature Extraction
András Kicsi, László Vidács, Viktor Csuvik, Ferenc Horváth, Árpád Beszédes, Ferenc Kocsis |
ICSR | 4 |
| 2016 | Negative Effects of Bytecode Instrumentation on Java Source Code CoverageabstractCode coverage measurement is an important element in white-box testing, both in industrial practice and academic research. Other related areas are highly dependent on code coverage as well, including test case generation, test prioritization, fault localization, and others. Inaccuracies of a code coverage tool sometimes do not matter that much but in certain situations they can lead to serious confusion. For Java, the prevalent approach to code coverage measurement is to use bytecode instrumentation due to its various benefits over source code instrumentation. However, if the results are to be mapped back to source code this may lead to inaccuracies due to the differences between the two program representations. In this paper, we systematically investigate the amount of differences in the results of these two Java code coverage approaches, enumerate the possible reasons and discuss the implications on various applications. For this purpose, we relied on two widely used tools to represent the two approaches and a set of benchmark programs from the open source domain. Dávid Tengeri, Ferenc Horváth, Árpád Beszédes, Tamás Gergely, Tibor Gyimóthy |
SANER | 2 |