Tamás Gergely

dblp:35/1518 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
2since 2021 · last 2025
0000-0001-7504-3580ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Weighted Call Frequency-Based Fault Localization
abstract
Spectrum-based fault localization is an automated technique that helps developers identify and isolate the origin of suspicious errors during software development. Despite being a well-researched topic, it is rarely used in the industry. The primary reason is that, in its basic version, it uses only local information on the coverage of a program element to estimate its probability of failure, rarely utilizing additional contextual information on the element or the test cases. Other researchers have tried solving the problem using contextual information with varying success. In this paper, we enhance the approach called Call Frequency-based Fault Localization, which analyzes the method's occurrence frequency in call-stack instances of failed tests. While it boosts SBFL's effectiveness, it overlooks the test scope. We propose that identifying unit and unit-like tests, followed by adjusting the frequency of the method by test type, can further enhance the fault localization ability of FL techniques. We empirically evaluated our method's effectiveness with the Defects4J benchmark. We found that utilizing weights in Call Frequency-based Fault Localization often ranks faulty methods higher, increasing the number of items in the top-10 positions.
Attila Szatmári, Aondowase James Orban, Tamás Gergely
ICST3
2023 A Case Against Coverage-Based Program Spectra
abstract
Spectrum-Based Fault Localization (SBFL) is a semi-automated debugging technique that gained popularity in the last decades due to its intuitive approach and relatively simple implementability. Despite this, the performance of practical SBFL techniques in terms of fault localization capability does not reach the threshold that would enable their acceptance by professional programmers. Almost all modern SBFL approaches are based on the code coverage-based spectrum, and on the assumption that a code element covered by failing tests should be treated as suspicious. However, it is easy to see that this is an over-approximation because many code elements may be executed that do not contribute to the test output, hence serving as noise in the process. A possible solution is to use backward dynamic program slices as program spectra computed from the output statement as the criterion, instead of the coverage. There are very few theoretical and practical results about this approach, so in this work we revisit the method and show how much more inferior coverage-based spectra are compared to slice-based spectra, both on theoretical and practical levels. We argue that code coverage-based SBFL is currently in a research pit due to this inherent approximation, and research on slice-based spectra should once more attain a much higher focus.
Péter Attila Soha, Tamás Gergely, Ferenc Horváth, Béla Vancsics, Árpád Beszédes
ICST2
2019 Understanding Test-to-Code Traceability Links: The Need for a Better Visualizing Model
Nadera Aljawabrah, Tamás Gergely, Mohammad Kharabsheh
ICCSA (4)2
2019 Differences between a static and a dynamic test-to-code traceability recovery method
abstract
Recovering test-to-code traceability links may be required in virtually every phase of development. This task might seem simple for unit tests thanks to two fundamental unit testing guidelines: isolation (unit tests should exercise only a single unit) and separation (they should be placed next to this unit). However, practice shows that recovery may be challenging because the guidelines typically cannot be fully followed. Furthermore, previous works have already demonstrated that fully automatic test-to-code traceability recovery for unit tests is virtually impossible in a general case. In this work, we propose a semi-automatic method for this task, which is based on computing traceability links using static and dynamic approaches, comparing their results and presenting the discrepancies to the user, who will determine the final traceability links based on the differences and contextual information. We define a set of discrepancy patterns, which can help the user in this task. Additional outcomes of analyzing the discrepancies are structural unit testing issues and related refactoring suggestions. For the static test-to-code traceability, we rely on the physical code structure, while for the dynamic, we use code coverage information. In both cases, we compute combined test and code clusters which represent sets of mutually traceable elements. We also present an empirical study of the method involving 8 non-trivial open source Java systems.
Tamás Gergely, Gergö Balogh, Ferenc Horváth, Béla Vancsics, Árpád Beszédes, Tibor Gyimóthy
Softw. Qual. J.1
2019 Code coverage differences of Java bytecode and source code instrumentation tools
Ferenc Horváth, Tamás Gergely, Árpád Beszédes, Dávid Tengeri, Gergö Balogh, Tibor Gyimóthy
Softw. Qual. J.2
2016 Are My Unit Tests in the Right Package?
abstract
The software development industry has adopted written and de facto standards for creating effective and maintainable unit tests. Unfortunately, like any other source code artifact, they are often written without conforming to these guidelines, or they may evolve into such a state. In this work, we address a specific type of issues related to unit tests. We seek to automatically uncover violations of two fundamental rules: 1) unit tests should exercise only the unit they were designed for, and 2) they should follow a clear packaging convention. Our approach is to use code coverage to investigate the dynamic behaviour of the tests with respect to the code elements of the program, and use this information to identify highly correlated groups of tests and code elements (using community detection algorithm). This grouping is then compared to the trivial grouping determined by package structure, and any discrepancies found are treated as "bad smells." We report on our related measurements on a set of large open source systems with notable unit test suites, and provide guidelines through examples for refactoring the problematic tests.
Gergö Balogh, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
SCAM2
2016 Negative Effects of Bytecode Instrumentation on Java Source Code Coverage
abstract
Code coverage measurement is an important element in white-box testing, both in industrial practice and academic research. Other related areas are highly dependent on code coverage as well, including test case generation, test prioritization, fault localization, and others. Inaccuracies of a code coverage tool sometimes do not matter that much but in certain situations they can lead to serious confusion. For Java, the prevalent approach to code coverage measurement is to use bytecode instrumentation due to its various benefits over source code instrumentation. However, if the results are to be mapped back to source code this may lead to inaccuracies due to the differences between the two program representations. In this paper, we systematically investigate the amount of differences in the results of these two Java code coverage approaches, enumerate the possible reasons and discuss the implications on various applications. For this purpose, we relied on two widely used tools to represent the two approaches and a set of benchmark programs from the open source domain.
Dávid Tengeri, Ferenc Horváth, Árpád Beszédes, Tamás Gergely, Tibor Gyimóthy
SANER4
2015 Empirical investigation of SEA-based dependence cluster properties
Árpád Beszédes, Lajos Schrettner, Béla Csaba, Tamás Gergely, Judit Jász, Tibor Gyimóthy
Sci. Comput. Program.4
2014 Impact analysis in the presence of dependence clusters using Static Execute After in WebKit
abstract
SUMMARY Impact analysis based on code dependence can provide opportunities to identify parts of the software affected by a change. Because changes usually have far reaching effects in programs, effective and efficient impact analysis is vital. Static Execute After (SEA) is a relation on procedures that is efficiently computable and accurate enough to be a candidate for the use in impact analysis in practice. To assess the applicability of SEA in terms of capturing real defects, we present results on integrating it into the build system of WebKit, a large, open source software system, and on related experiments. We show that a large number of real defects can be captured by impact sets computed by SEA, albeit many of them are large. We demonstrate that this is not an issue in applying it to regression test prioritization, but generally it can be an obstacle in the path to efficient use of impact analysis. We believe that the main reason for large impact sets is the formation of dependence clusters in code. As apparently dependence clusters cannot be easily avoided in the majority of cases, we focus on determining the effects these clusters have on impact analysis and regression test prioritization. Copyright © 2013 John Wiley & Sons, Ltd.
Lajos Schrettner, Judit Jász, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
J. Softw. Evol. Process.3
2013 Empirical investigation of SEA-based dependence cluster properties
abstract
Dependence clusters are (maximal) groups of source code entities that each depend on the other according to some dependence relation. Such clusters are generally seen as detrimental to many software engineering activities, but their formation and overall structure are not well understood yet. In a set of subject programs from moderate to large sizes, we observed frequent occurrence of dependence clusters using Static Execute After (SEA) dependences (SEA is a conservative yet efficiently computable dependence relation on program procedures). We identified potential linchpins inside the clusters; these are procedures that can primarily be made responsible for keeping the cluster together. Furthermore, we found that as the size of the system increases, it is more likely that multiple procedures are jointly responsible as sets of linchpins. We also give a heuristic method based on structural metrics for locating possible linchpins as their exact identification is unfeasible in practice, and presently there are no better ways than the brute-force method. We defined novel metrics and comparison methods to be able to demonstrate clusters of different sizes in programs.
Árpád Beszédes, Lajos Schrettner, Béla Csaba, Tamás Gergely, Judit Jász, Tibor Gyimóthy
SCAM4
2012 Code coverage-based regression test selection and prioritization in WebKit
abstract
Automated regression testing is often crucial in order to maintain the quality of a continuously evolving software system. However, in many cases regression test suites tend to grow too large to be suitable for full re-execution at each change of the software. In this case selective retesting can be applied to reduce the testing cost while maintaining similar defect detection capability. One of the basic test selection methods is the one based on code coverage information, where only those tests are included that cover some parts of the changes. We experimentally applied this method to the open source web browser engine project WebKit to find out the technical difficulties and the expected benefits if this method is to be introduced into the actual build process. Although the principle is simple, we had to solve a number of technical issues, so we report how this method was adapted to be used in the official build environment. Second, we present results about the selection capabilities for a selected set of revisions of WebKit, which are promising. We also applied different test case prioritization strategies to further reduce the number of tests to execute. We explain these strategies and compare their usefulness in terms of defect detection and test suite reduction.
Árpád Beszédes, Tamás Gergely, Lajos Schrettner, Judit Jász, Laszlo Lango, Tibor Gyimóthy
ICSM2
2012 Impact Analysis in the Presence of Dependence Clusters Using Static Execute after in WebKit
abstract
Impact analysis based on code dependence can be an integral part of software quality assurance by providing opportunities to identify those parts of the software system that are affected by a change. Because changes usually have far reaching effects in programs, effective and efficient impact analysis is vital, which has different applications including change propagation and regression testing. Static Execute After (SEA) is a relation on program elements (procedures) that is efficiently computable and accurate enough to be a candidate for use in impact analysis in practice. To assess the applicability of SEA in terms of capturing real defects, we present results on integrating it into the build system of Web Kit, a large, open source software system, and on related experiments. We show that a large number of real defects can be captured by impact sets computed by SEA, albeit many of them are large. We demonstrate that this is not an issue in applying it to regression test prioritization, but generally it can be an obstacle in the path to efficient use of impact analysis. We believe that the main reason for large impact sets is the formation of dependence clusters in code. As apparently dependence clusters cannot be easily avoided in the majority of cases, we focus on determining the effects these clusters have on impact analysis.
Lajos Schrettner, Judit Jász, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
SCAM3
2010 Effect of test completeness and redundancy measurement on post release failures - An industrial experience report
abstract
In risk-based testing, compromises are often made to release a system in spite of knowing that it has outstanding defects. In an industrial setting, time and cost are often the “exit criteria” and - unfortunately - not the technical aspects like coverage or defect ratio. In such situations, the stakeholders accept that the remaining defects will be found after release, so sufficient resources are allocated to the “stabilization” phases following the release. It is hard for many organizations to see that such an approach is significantly costlier than trying to locate the defects earlier. We performed an empirical investigation of this for one of our industrial partners (a financial company). In this project, significant perfective maintenance was performed on the large information system. Based on changes made to the system, we carried out procedure level code coverage measurements with code level change impact analysis, and a similarity-based comparison of test cases in order to quantitatively check the completeness and redundancy of the tests performed. In addition, we logged and compared the number of defects found during testing and live operation. The data obtained were surprising for both the developers and the customer as well, leading to a major reorganization of their development, testing, and operation processes. After the reorganization, a significant improvement in these indicators for testing efficiency was observed.
Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy, Milan Imre Gyalai
ICSM1
2007 Computation of Static Execute After Relation with Applications to Software Maintenance
abstract
In this paper, we introduce static execute after (SEA) relationship among program components and present an efficient analysis algorithm. Our case studies show that SEA may approximate static slicing with perfect recall and high precision, while being much less expensive and more usable. When differentiating between explicit and hidden dependencies, our case studies also show that SEA may correlate with direct and indirect class coupling. We speculate that SEA may find applications in computation of hidden dependencies and through it in many maintenance tasks, including change propagation and regression testing.
Árpád Beszédes, Tamás Gergely, Judit Jász, Gabriella Tóth, Tibor Gyimóthy, Václav Rajlich
ICSM2
2002 Multi-access Services for the Management of Diabetes Mellitus: The M2DM Project
Riccardo Bellazzi, Giuliana Bensa, Eulalia Brugués, Ewart R. Carson, Claudio Cobelli, Derek G. Cramp, Giuseppe d'Annunzio, Pasquale De Cata, Alberto de Leiva, Tibor Deutsch, Pietro Fratino, Carmine Gazzaruso, Angel Garcia, Tamás Gergely, Enrique J. Gómez, Fiona E. Harvey, Pietro Ferrari, Christiane Harras Friederich, María Elena Hernando, Maged N. Kamel Boulos, Cristiana Larizza, Hans Ludekke, Monika Luebker, Alberto Maran, Gianluca Nucci, Fernando Ortiz Garcia, Cristina Pennati, Abdul V. Roudsari, Mercedes Rigla, Karsten Schutte, Mario Stefanelli
AMIA14
1987 Inductive Inference on the Base of Fixed Point Theory
Tamás Gergely
IJCAI1
1983 Negative Hyper-Resolution for Proving Statements Containing Transitive Relations
Tamás Gergely, Konstantin Vershinin
IJCAI1
1982 A Theory of Interactive Programming
Tamás Gergely, László Úry
Acta Informatica1
1980 Model Theoretic Semantics For Many-Purpose Languages And Language Hierarchies
Hajnal Andréka, Tamás Gergely, István Németi
COLING2
1975 On the Role of Mathematical Language Concept in the Theory of Intelligent Systems
Hajnal Andréka, Tamás Gergely, István Németi
IJCAI2