Tibor Gyimóthy

dblp:23/234 · DBLP profile ↗
← Back
74ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-2123-7387ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 60 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7Artificial intelligence and machine learning · 4Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2024 Context Switch Sensitive Fault Localization
abstract
Spectrum-Based Fault Localization (SBFL) is a popular technique to assist developers in pinpointing faulty elements within their code based on test outcomes and code coverage. In this paper, we examine the impact of context switching, i.e., when developers must frequently shift their attention between different code parts (such as methods and classes) while going down the SBFL ranked list to find the faulty statement. The basis of our study is the observation that it requires less effort to investigate statements that are next to each other rather than those in different methods and classes. In particular, we analyse the number of visited methods and classes, as well as the frequency of switches between them during the fault localization process. We found that, in programs from the Defects4J benchmark, developers need to explore 40 methods and 12 classes on average, before finding the faulty statement, leading to 53 method- and 40 class switches, respectively.
Ferenc Horváth, Roland Aszmann, Péter Attila Soha, Árpád Beszédes, Tibor Gyimóthy
EASE5
2024 On the Stability and Applicability of Deep Learning in Fault Localization
abstract
Numerous Deep Learning (DL)-based fault localization (FL) methods are developed with the aim of leveraging the code coverage matrix and failure vector to identify the connection between program elements and defects. The imbalanced data on which these approaches train their models poses a substantial challenge to the effectiveness of fault localization techniques. This study explores the stability of fault localization models in deep learning, specifically, their performance when trained repeatedly using the same input but varying random initializations. Using the Defect4J benchmark, we trained deep learning models (MLP, CNN, and RNN) independently and found that 86 cases resulted in (partly) consistent rankings among all five models and versions, while 621 exhibited varying outcomes, meaning that 90 % of the produced ranks were different in subsequent trainings. The models showed significant variability in ranking results, with maximum ranks sometimes five times that of the minimum. We also adapted the churn metric from DL research to evaluate models, confirming their instability. To improve stability, meta-parameter optimization, model simplification and resampling has been applied. Although some of these techniques proved effective, even with the improvements, the models remained insufficiently stable to produce reliable results.
Viktor Csuvik, Roland Aszmann, Árpád Beszédes, Ferenc Horváth, Tibor Gyimóthy
SANER5
2023 Can ChatGPT Fix My Code?
Viktor Csuvik, Tibor Gyimóthy, László Vidács
ICSOFT2
2022 Using contextual knowledge in interactive fault localization
abstract
Abstract Tool support for automated fault localization in program debugging is limited because state-of-the-art algorithms often fail to provide efficient help to the user. They usually offer a ranked list of suspicious code elements, but the fault is not guaranteed to be found among the highest ranks. In Spectrum-Based Fault Localization (SBFL) – which uses code coverage information of test cases and their execution outcomes to calculate the ranks –, the developer has to investigate several locations before finding the faulty code element. Yet, all the knowledge she a priori has or acquires during this process is not reused by the SBFL tool. There are existing approaches in which the developer interacts with the SBFL algorithm by giving feedback on the elements of the prioritized list. We propose a new approach called iFL which extends interactive approaches by exploiting contextual knowledge of the user about the next item in the ranked list (e. g., a statement), with which larger code entities (e. g., a whole function) can be repositioned in their suspiciousness. We implemented a closely related algorithm proposed by Gong et al., called Talk. First, we evaluated iFL using simulated users, and compared the results to SBFL and Talk. Next, we introduced two types of imperfections in the simulation: user’s knowledge and confidence levels. On SIR and Defects4J, results showed notable improvements in fault localization efficiency, even with strong user imperfections. We then empirically evaluated the effectiveness of the approach with real users in two sets of experiments: a quantitative evaluation of the successfulness of using iFL, and a qualitative evaluation of practical uses of the approach with experienced developers in think-aloud sessions.
Ferenc Horváth, Árpád Beszédes, Béla Vancsics, Gergö Balogh, László Vidács, Tibor Gyimóthy
Empir. Softw. Eng.6
2020 Experiments with Interactive Fault Localization Using Simulated and Real Users
abstract
Fault localization is considered a difficult and time consuming activity. However, tool support for automated fault localization is still limited because state-of-the-art algorithms often fail to provide efficient help to the user. They usually offer a ranked list of suspicious code elements, but the fault is not guaranteed to be found among the highest ranks. In Spectrum-Based Fault Localization (SBFL) - which uses code coverage information of test cases and their execution outcomes to calculate the ranks -, the developer has to investigate several locations before finding the faulty code element. Yet, all the knowledge she a priori has or acquires during this process is not reused by the SBFL tool. We propose an approach in which the developer interacts with the SBFL algorithm by giving feedback on the elements of the prioritized list. We exploit contextual knowledge of the user about the next item in the ranked list (e. g., a statement), with which larger code entities (e. g., a whole function) can be repositioned in their suspiciousness. First, we evaluated the approach using simulated users incorporating two types of imperfections, their knowledge and confidence levels. On SIR and Defects4J, results showed notable improvements in fault localization efficiency, even with strong user imperfections. We then empirically evaluated the effectiveness of the approach with real users, which also showed promising results.
Ferenc Horváth, Árpád Beszédes, Béla Vancsics, Gergö Balogh, László Vidács, Tibor Gyimóthy
ICSME6
2020 TestRoutes: A Manually Curated Method Level Dataset for Test-to-Code Traceability
abstract
High test-to-code traceability can be an important aspect of quality assurance and can contribute to bug localization and code maintenance. Several existing techniques and a considerable effort from the scientific community already made significant advances in the field. Despite this, readily accessible data on traceability links is very scarce. To contribute to related research, we present a manually curated test-to-code traceability dataset containing the traceability information on 220 test cases. This method-level data was gathered from 4 open-source software systems written in the Java language, distinguishing not only focal information on test cases but also highlighting the utilized helper methods on both the test and production aspects of code. The data includes more than 2000 of such method classifications.
András Kicsi, László Vidács, Tibor Gyimóthy
MSR3
2020 Leveraging Contextual Information from Function Call Chains to Improve Fault Localization
abstract
In Spectrum Based Fault Localization, program elements such as statements or functions are ranked according to a suspiciousness score which can guide the programmer in finding the fault more efficiently. However, such a ranking does not include any additional information about the suspicious code elements. In this work, we propose to complement function-level spectrum based fault localization with function call chains - i.e., snapshots of the call stack occurring during execution - on which the fault localization is first performed, and then narrowed down to functions. Our experiments using defects from Defects4J show that (i) 69% of the defective functions can be found in call chains with highest scores, (ii) in 4 out of 6 cases the proposed approach can improve Ochiai ranking of 1 to 9 positions on average, with a relative improvement of 19–48%, and (iii) the improvement is substantial (66–98%) when Ochiai produces bad rankings for the faulty functions.
Árpád Beszédes, Ferenc Horváth, Massimiliano Di Penta, Tibor Gyimóthy
SANER4
2020 An automatically created novel bug dataset and its validation in bug prediction
abstract
Bugs are inescapable during software development due to frequent code changes, tight deadlines, etc.; therefore, it is important to have tools to find these errors. One way of performing bug identification is to analyze the characteristics of buggy source code elements from the past and predict the present ones based on the same characteristics, using e.g. machine learning models. To support model building tasks, code elements and their characteristics are collected in so-called bug datasets which serve as the input for learning. We present the BugHunter Dataset: a novel kind of automatically constructed and freely available bug dataset containing code elements (files, classes, methods) with a wide set of code metrics and bug information. Other available bug datasets follow the traditional approach of gathering the characteristics of all source code elements (buggy and non-buggy) at only one or more pre-selected release versions of the code. Our approach, on the other hand, captures the buggy and the fixed states of the same source code elements from the narrowest timeframe we can identify for a bug’s presence, regardless of release versions. To show the usefulness of the new dataset, we built and evaluated bug prediction models and achieved F-measure values over 0.74.
Rudolf Ferenc, Péter Gyimesi, Gábor Gyimesi, Zoltán Tóth, Tibor Gyimóthy
J. Syst. Softw.5
2020 A public unified bug dataset for java and its assessment regarding metrics and bug prediction
abstract
Abstract Bug datasets have been created and used by many researchers to build and validate novel bug prediction models. In this work, our aim is to collect existing public source code metric-based bug datasets and unify their contents. Furthermore, we wish to assess the plethora of collected metrics and the capabilities of the unified bug dataset in bug prediction. We considered 5 public datasets and we downloaded the corresponding source code for each system in the datasets and performed source code analysis to obtain a common set of source code metrics. This way, we produced a unified bug dataset at class and file level as well. We investigated the diversion of metric definitions and values of the different bug datasets. Finally, we used a decision tree algorithm to show the capabilities of the dataset in bug prediction. We found that there are statistically significant differences in the values of the original and the newly calculated metrics; furthermore, notations and definitions can severely differ. We compared the bug prediction capabilities of the original and the extended metric suites (within-project learning). Afterwards, we merged all classes (and files) into one large dataset which consists of 47,618 elements (43,744 for files) and we evaluated the bug prediction model build on this large dataset as well. Finally, we also investigated cross-project capabilities of the bug prediction models and datasets. We made the unified dataset publicly available for everyone. By using a public unified dataset as an input for different bug prediction related investigations, researchers can make their studies reproducible, thus able to be validated and verified.
Rudolf Ferenc, Zoltán Tóth, Gergely Ladányi, István Siket, Tibor Gyimóthy
Softw. Qual. J.5
2019 Towards an Accurate Prediction of the Question Quality on Stack Overflow using a Deep-Learning-Based NLP Approach
László Tóth 0002, Balázs Nagy 0004, Dávid Janthó, László Vidács, Tibor Gyimóthy
ICSOFT5
2019 A Mobile IoT Device Simulator for IoT-Fog-Cloud Systems
Attila Kertész, Tamas Pflanzner, Tibor Gyimóthy
J. Grid Comput.3
2019 Feature analysis using information retrieval, community detection and structural analysis methods in product line adoption
abstract
In industrial practice the clone-and-own strategy is often applied when in the pressure of high demand of customized features. The adoption of software product line (SPL) architecture is a large one time investment that affects both technical and organizational issues. The analysis of the feature structure is a crucial point in the SPL adoption process involving domain experts working at a higher level of abstraction and developers working directly on the program code. We propose automatic methods to extract feature-to-program links starting from very high level set of features provided by domain experts. For this purpose we combine call graph information with textual similarity between code and high level features. In addition, in depth understanding of the feature structure is supported by finding communities between programs and relating them to features. As features are originated from domain experts, community analysis reveals discrepancies between expert view and internal code structure. We found that communities correspond well to the high level features, with usually more than half of feature code located in specialized communities. We report experiments at two levels of features and more than 2000 Magic 4GL programs in an industrial SPL adoption project.
András Kicsi, Viktor Csuvik, László Vidács, Ferenc Horváth, Árpád Beszédes, Tibor Gyimóthy, Ferenc Kocsis
J. Syst. Softw.6
2019 Differences between a static and a dynamic test-to-code traceability recovery method
abstract
Recovering test-to-code traceability links may be required in virtually every phase of development. This task might seem simple for unit tests thanks to two fundamental unit testing guidelines: isolation (unit tests should exercise only a single unit) and separation (they should be placed next to this unit). However, practice shows that recovery may be challenging because the guidelines typically cannot be fully followed. Furthermore, previous works have already demonstrated that fully automatic test-to-code traceability recovery for unit tests is virtually impossible in a general case. In this work, we propose a semi-automatic method for this task, which is based on computing traceability links using static and dynamic approaches, comparing their results and presenting the discrepancies to the user, who will determine the final traceability links based on the differences and contextual information. We define a set of discrepancy patterns, which can help the user in this task. Additional outcomes of analyzing the discrepancies are structural unit testing issues and related refactoring suggestions. For the static test-to-code traceability, we rely on the physical code structure, while for the dynamic, we use code coverage information. In both cases, we compute combined test and code clusters which represent sets of mutually traceable elements. We also present an empirical study of the method involving 8 non-trivial open source Java systems.
Tamás Gergely, Gergö Balogh, Ferenc Horváth, Béla Vancsics, Árpád Beszédes, Tibor Gyimóthy
Softw. Qual. J.6
2019 Code coverage differences of Java bytecode and source code instrumentation tools
Ferenc Horváth, Tamás Gergely, Árpád Beszédes, Dávid Tengeri, Gergö Balogh, Tibor Gyimóthy
Softw. Qual. J.6
2019 Prediction models for performance, power, and energy efficiency of software executed on heterogeneous hardware
Dénes Bán, Rudolf Ferenc, István Siket, Ákos Kiss 0001, Tibor Gyimóthy
J. Supercomput.5
2018 Feature Level Complexity and Coupling Analysis in 4GL Systems
András Kicsi, Viktor Csuvik, László Vidács, Árpád Beszédes, Tibor Gyimóthy
ICCSA (5)5
2018 [Research Paper] Static JavaScript Call Graphs: A Comparative Study
abstract
The popularity and wide adoption of JavaScript both at the client and server side makes its code analysis more important than ever before. Most of the algorithms for vulnerability analysis, coding issue detection, or type inference rely on the call graph representation of the underlying program. Despite some obvious advantages of dynamic analysis, static algorithms should also be considered for call graph construction as they do not require extensive test beds for programs and their costly execution and tracing. In this paper, we systematically compare five widely adopted static algorithms - implemented by the npm call graph, IBM WALA, Google Closure Compiler, Approximate Call Graph, and Type Analyzer for JavaScript tools - for building JavaScript call graphs on 26 WebKit SunSpider benchmark programs and 6 real-world Node.js modules. We provide a performance analysis as well as a quantitative and qualitative evaluation of the results. We found that there was a relatively large intersection of the found call edges among the algorithms, which proved to be 100% precise. However, most of the tools found edges that were missed by all others. ACG had the highest precision followed immediately by TAJS, but ACG found significantly more call edges. As for the combination of tools, ACG and TAJS together covered 99% of the found true edges by all algorithms, while maintaining a precision as high as 98%. Only two of the tools were able to analyze up-to-date multi-file Node.js modules due to incomplete language features support. They agreed on almost 60% of the call edges, but each of them found valid edges that the other missed.
Gabor Antal, Péter Hegedüs, Zoltán Tóth, Rudolf Ferenc, Tibor Gyimóthy
SCAM5
2018 Empirical evaluation of software maintainability based on a manually validated refactoring dataset
Péter Hegedüs, István Kádár, Rudolf Ferenc, Tibor Gyimóthy
Inf. Softw. Technol.4
2017 Coarse Hierarchical Delta Debugging
abstract
This paper introduces the Coarse Hierarchical Delta Debugging algorithm for efficient test case reduction. It can be used as a test case simplification algorithm in its own right if theoretical minimality is not a strict requirement, or it can act as a preprocessing step to the original Hierarchical Delta Debugging algorithm. Evaluation of artificial and real test cases shows that a coarse variant can produce reduced test cases with significantly fewer testing steps than the original algorithm (58% gain on average, 79% maximum), while still keeping the outputs acceptably small (never increasing the reduced test cases by more than 0.36% of the input).
Renáta Hodován, Ákos Kiss 0001, Tibor Gyimóthy
ICSME3
2017 Empirical study on refactoring large-scale industrial systems and its effects on maintainability
Gábor Szoke, Gabor Antal, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy
J. Syst. Softw.5
2016 Assessment of the Code Refactoring Dataset Regarding the Maintainability of Methods
István Kádár, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy
ICCSA (4)4
2016 Are My Unit Tests in the Right Package?
abstract
The software development industry has adopted written and de facto standards for creating effective and maintainable unit tests. Unfortunately, like any other source code artifact, they are often written without conforming to these guidelines, or they may evolve into such a state. In this work, we address a specific type of issues related to unit tests. We seek to automatically uncover violations of two fundamental rules: 1) unit tests should exercise only the unit they were designed for, and 2) they should follow a clear packaging convention. Our approach is to use code coverage to investigate the dynamic behaviour of the tests with respect to the code elements of the program, and use this information to identify highly correlated groups of tests and code elements (using community detection algorithm). This grouping is then compared to the trivial grouping determined by package structure, and any discrepancies found are treated as "bad smells." We report on our related measurements on a set of large open source systems with notable unit test suites, and provide guidelines through examples for refactoring the problematic tests.
Gergö Balogh, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
SCAM4
2016 A Code Refactoring Dataset and Its Assessment Regarding Software Maintainability
abstract
It is very common in various fields that there is a gap between theoretical results and their practical applications. This is true for code refactoring as well, which has a solid theoretical background while being used in development practice at the same time. However, more and more studies suggest that developers perform code refactoring entirely differently than the theory would suggest. Our paper encourages the further investigation of code refactorings in practice by providing an excessive open dataset of source code metrics and applied refactorings through several releases of 7 open-source systems. As a first step of processing this dataset, we examined the quality attributes of the refactored source code classes and the values of source code metrics improved by those refactorings. Our early results show that lower maintainability indeed triggers more code refactorings in practice and these refactorings significantly decrease complexity, code lines, coupling and clone metrics. However, we observed a decrease in comment related metrics in the refactored code.
István Kádár, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy
SANER4
2016 Designing and Developing Automated Refactoring Transformations: An Experience Report
abstract
There are several challenges which should be kept in mind during the design and development phases of a refactoring tool, and one is that developers have several expectations that are quite hard to satisfy. In this report, we present our experiences of a two-year project where we attempted to create an automatic refactoring tool. In this project, we worked with five software development companies that wanted to improve the maintainability of their products. The project was designed to take into account the expectations of the developers of these companies and consisted of three main stages: a manual refactoring phase, a tool building phase, and an automatic refactoring phase. Throughout these stages we collected the opinions of the developers and faced several challenges on how to automate refactoring transformations, which we present and summarize.
Gábor Szoke, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy
SANER4
2016 Negative Effects of Bytecode Instrumentation on Java Source Code Coverage
abstract
Code coverage measurement is an important element in white-box testing, both in industrial practice and academic research. Other related areas are highly dependent on code coverage as well, including test case generation, test prioritization, fault localization, and others. Inaccuracies of a code coverage tool sometimes do not matter that much but in certain situations they can lead to serious confusion. For Java, the prevalent approach to code coverage measurement is to use bytecode instrumentation due to its various benefits over source code instrumentation. However, if the results are to be mapped back to source code this may lead to inaccuracies due to the differences between the two program representations. In this paper, we systematically investigate the amount of differences in the results of these two Java code coverage approaches, enumerate the possible reasons and discuss the implications on various applications. For this purpose, we relied on two widely used tools to represent the two approaches and a set of benchmark programs from the open source domain.
Dávid Tengeri, Ferenc Horváth, Árpád Beszédes, Tamás Gergely, Tibor Gyimóthy
SANER5
2015 Identifying wasted effort in the field via developer interaction data
abstract
During software projects, several parts of the source code are usually re-written due to imperfect solutions before the code is released. This wasted effort is of central interest to the project management to assure on-time delivery. Although the amount of thrown-away code can be measured from version control systems, stakeholders are more interested in productivity dynamics that reflect the constant change in a software project. In this paper we present a field study of measuring the productivity of a medium-sized J2EE project. We propose a productivity analysis method where productivity is expressed through dynamic profiles - the so-called Micro-Productivity Profiles (MPPs). They can be used to characterize various constituents of software projects such as components, phases and teams. We collected detailed traces of developers' actions using an Eclipse IDE plug-in for seven months of software development throughout two milestones. We present and evaluate profiles of two important axes of the development process: by milestone and by application layers. MPPs can be an aid to take project control actions and help in planning future projects. Based on the experiments, project stakeholders identified several points to improve the development process. It is also acknowledged, that profiles show additional information compared to a naive diff-based approach.
Gergö Balogh, Gabor Antal, Árpád Beszédes, László Vidács, Tibor Gyimóthy, Ádám Zoltán Végh
ICSME5
2015 Do automatic refactorings improve maintainability? An industrial case study
abstract
Refactoring is often treated as the main remedy against the unavoidable code erosion happening during software evolution. Studies show that refactoring is indeed an elemental part of the developers' arsenal. However, empirical studies about the impact of refactorings on software maintainability still did not reach a consensus. Moreover, most of these empirical investigations are carried out on open-source projects where distinguishing refactoring operations from other development activities is a challenge in itself. We had a chance to work together with several software development companies in a project where they got extra budget to improve their source code by performing refactoring operations. Taking advantage of this controlled environment, we collected a large amount of data during a refactoring phase where the developers used a (semi)automatic refactoring tool. By measuring the maintainability of the involved subject systems before and after the refactorings, we got valuable insights into the effect of these refactorings on large-scale industrial projects. All but one company, who applied a special refactoring strategy, achieved a maintainability improvement at the end of the refactoring phase, but even that one company suffered from the negative impact of only one type of refactoring.
Gábor Szoke, Csaba Nagy 0001, Péter Hegedüs, Rudolf Ferenc, Tibor Gyimóthy
ICSME5
2015 FaultBuster: An automatic code smell refactoring toolset
abstract
One solution to prevent the quality erosion of a software product is to maintain its quality by continuous refac-toring. However, refactoring is not always easy. Developers need to identify the piece of code that should be improved and decide how to rewrite it. Furthermore, refactoring can also be risky; that is, the modified code needs to be re-tested, so developers can see if they broke something. Many IDEs offer a range of refactorings to support so-called automatic refactoring, but tools which are really able to automatically refactor code smells are still under research. In this paper we introduce FaultBuster, a refactoring toolset which is able to support automatic refactoring: identifying the problematic code parts via static code analysis, running automatic algorithms to fix selected code smells, and executing integrated testing tools. In the heart of the toolset lies a refactoring framework to control the analysis and the execution of automatic algorithms. FaultBuster provides IDE plugins to interact with developers via popular IDEs (Eclipse, Netbeans and IntelliJ IDEA). All the tools were developed and tested in a 2-year project with 6 software development companies where thousands of code smells were identified and fixed in 5 systems having altogether over 5 million lines of code.
Gábor Szoke, Csaba Nagy 0001, Lajos Jeno Fülöp, Rudolf Ferenc, Tibor Gyimóthy
SCAM5
2015 Empirical investigation of SEA-based dependence cluster properties
Árpád Beszédes, Lajos Schrettner, Béla Csaba, Tamás Gergely, Judit Jász, Tibor Gyimóthy
Sci. Comput. Program.6
2014 Service Layer for IDE Integration of C/C++ Preprocessor Related Analysis
Richárd Dévai, László Vidács, Rudolf Ferenc, Tibor Gyimóthy
ICCSA (5)4
2014 A Case Study of Refactoring Large-Scale Industrial Systems to Efficiently Improve Source Code Quality
Gábor Szoke, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy
ICCSA (5)4
2014 Source Meter Sonar Qube Plug-in
abstract
The SourceMeter Sonar Qube plug-in is an extension of Sonar Qube, an open-source platform for managing code quality made by Sonar Source S.A, Switzerland. The plug-in extends the built-in Java code analysis engine of Sonar Qube with Front End ART's high-end Java code analysis engine. Most of Sonar Qubes original analysis results are replaced (including the detected source code duplications), while the range of available analyses is extended with a number of additional metrics and issue detectors. Additionally, the plug-in offers new GUI features on the Sonar Qube dashboard and drilldown views, making the Sonar Qube user experience more comfortable and the work with the tool more productive.
Rudolf Ferenc, Laszlo Lango, István Siket, Tibor Gyimóthy, Tibor Bakota
SCAM4
2014 Bulk Fixing Coding Issues and Its Effects on Software Quality: Is It Worth Refactoring?
abstract
The quality of a software system is mostly defined by its source code. Software evolves continuously, it gets modified, enhanced, and new requirements always arise. If we do not spend time periodically on improving our source code, it becomes messy and its quality will decrease inevitably. Literature tells us that we can improve the quality of our software product by regularly refactoring it. But does refactoring really increase software quality? Can it happen that a refactoring decreases the quality? Is it possible to recognize the change in quality caused by a single refactoring operation? In our paper, we seek answers to these questions in a case study of refactoring large-scale proprietary software systems. We analyzed the source code of 5 systems, and measured the quality of several revisions for a period of time. We analyzed 2 million lines of code and identified nearly 200 refactoring commits which fixed over 500 coding issues. We found that one single refactoring only makes a small change (sometimes even decreases quality), but when we do them in blocks, we can significantly increase quality, which can result not only in the local, but also in the global improvement of the code.
Gábor Szoke, Gabor Antal, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy
SCAM5
2014 Toolset and Program Repository for Code Coverage-Based Test Suite Analysis and Manipulation
abstract
Code coverage is often used in academic and industrial practice of white-box software testing. Various test optimization methods, e.g. Test selection and prioritization, rely on code coverage information, but other related fields benefit from it as well, such as fault localization. These methods require access to the fine details of coverage information and efficient ways of processing this data. The purpose of the (free) SoDA library and toolset is to provide an efficient set of data structures and algorithms which can be used to prepare, store and analyze in various ways data related to code coverage. The focus of SoDA is not on the calculation of coverage data (such as instrumentation and test execution) but on the analysis and manipulation of test suites based on such information. An important design goal of the library was to be usable on industrial-size programs and test suites. Furthermore, there is no limitation on programming language, analysis granularity and coverage criteria. In this paper, we demonstrate the purpose and benefits of the library, the associated toolset, which also includes a graphical user interface, as well as possible usage scenarios. SoDA also includes a repository of prepared programs, which are from small to large sizes and can be used for experimentation and as a benchmark for code coverage related research.
Dávid Tengeri, Árpád Beszédes, David Havas, Tibor Gyimóthy
SCAM4
2014 Impact analysis in the presence of dependence clusters using Static Execute After in WebKit
abstract
SUMMARY Impact analysis based on code dependence can provide opportunities to identify parts of the software affected by a change. Because changes usually have far reaching effects in programs, effective and efficient impact analysis is vital. Static Execute After (SEA) is a relation on procedures that is efficiently computable and accurate enough to be a candidate for the use in impact analysis in practice. To assess the applicability of SEA in terms of capturing real defects, we present results on integrating it into the build system of WebKit, a large, open source software system, and on related experiments. We show that a large number of real defects can be captured by impact sets computed by SEA, albeit many of them are large. We demonstrate that this is not an issue in applying it to regression test prioritization, but generally it can be an obstacle in the path to efficient use of impact analysis. We believe that the main reason for large impact sets is the formation of dependence clusters in code. As apparently dependence clusters cannot be easily avoided in the majority of cases, we focus on determining the effects these clusters have on impact analysis and regression test prioritization. Copyright © 2013 John Wiley & Sons, Ltd.
Lajos Schrettner, Judit Jász, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
J. Softw. Evol. Process.5
2013 A Methodology and Framework for Automatic Layout Independent GUI Testing of Applications Developed in Magic xpa
Daniel Fritsi, Csaba Nagy 0001, Rudolf Ferenc, Tibor Gyimóthy
ICCSA (2)4
2013 Empirical investigation of SEA-based dependence cluster properties
abstract
Dependence clusters are (maximal) groups of source code entities that each depend on the other according to some dependence relation. Such clusters are generally seen as detrimental to many software engineering activities, but their formation and overall structure are not well understood yet. In a set of subject programs from moderate to large sizes, we observed frequent occurrence of dependence clusters using Static Execute After (SEA) dependences (SEA is a conservative yet efficiently computable dependence relation on program procedures). We identified potential linchpins inside the clusters; these are procedures that can primarily be made responsible for keeping the cluster together. Furthermore, we found that as the size of the system increases, it is more likely that multiple procedures are jointly responsible as sets of linchpins. We also give a heuristic method based on structural metrics for locating possible linchpins as their exact identification is unfeasible in practice, and presently there are no better ways than the brute-force method. We defined novel metrics and comparison methods to be able to demonstrate clusters of different sizes in programs.
Árpád Beszédes, Lajos Schrettner, Béla Csaba, Tamás Gergely, Judit Jász, Tibor Gyimóthy
SCAM6
2012 A cost model based on software maintainability
abstract
In this paper we present a maintainability based model for estimating the costs of developing source code in its evolution phase. Our model adopts the concept of entropy in thermodynamics, which is used to measure the disorder of a system. In our model, we use maintainability for measuring disorder (i.e. entropy) of the source code of a software system. We evaluated our model on three proprietary and two open source real world software systems implemented in Java, and found that the maintainability of these evolving software is decreasing over time. Furthermore, maintainability and development costs are in exponential relationship with each other. We also found that our model is able to predict future development costs with high accuracy in these systems.
Tibor Bakota, Péter Hegedüs, Gergely Ladányi, Peter Kortvelyesi, Rudolf Ferenc, Tibor Gyimóthy
ICSM6
2012 Code coverage-based regression test selection and prioritization in WebKit
abstract
Automated regression testing is often crucial in order to maintain the quality of a continuously evolving software system. However, in many cases regression test suites tend to grow too large to be suitable for full re-execution at each change of the software. In this case selective retesting can be applied to reduce the testing cost while maintaining similar defect detection capability. One of the basic test selection methods is the one based on code coverage information, where only those tests are included that cover some parts of the changes. We experimentally applied this method to the open source web browser engine project WebKit to find out the technical difficulties and the expected benefits if this method is to be introduced into the actual build process. Although the principle is simple, we had to solve a number of technical issues, so we report how this method was adapted to be used in the official build environment. Second, we present results about the selection capabilities for a selected set of revisions of WebKit, which are promising. We also applied different test case prioritization strategies to further reduce the number of tests to execute. We explain these strategies and compare their usefulness in terms of defect detection and test suite reduction.
Árpád Beszédes, Tamás Gergely, Lajos Schrettner, Judit Jász, Laszlo Lango, Tibor Gyimóthy
ICSM6
2012 Impact Analysis in the Presence of Dependence Clusters Using Static Execute after in WebKit
abstract
Impact analysis based on code dependence can be an integral part of software quality assurance by providing opportunities to identify those parts of the software system that are affected by a change. Because changes usually have far reaching effects in programs, effective and efficient impact analysis is vital, which has different applications including change propagation and regression testing. Static Execute After (SEA) is a relation on program elements (procedures) that is efficiently computable and accurate enough to be a candidate for use in impact analysis in practice. To assess the applicability of SEA in terms of capturing real defects, we present results on integrating it into the build system of Web Kit, a large, open source software system, and on related experiments. We show that a large number of real defects can be captured by impact sets computed by SEA, albeit many of them are large. We demonstrate that this is not an issue in applying it to regression test prioritization, but generally it can be an obstacle in the path to efficient use of impact analysis. We believe that the main reason for large impact sets is the formation of dependence clusters in code. As apparently dependence clusters cannot be easily avoided in the majority of cases, we focus on determining the effects these clusters have on impact analysis.
Lajos Schrettner, Judit Jász, Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy
SCAM5
2011 Complexity Measures in 4GL Environment
Csaba Nagy 0001, László Vidács, Rudolf Ferenc, Tibor Gyimóthy, Ferenc Kocsis
ICCSA (5)4
2011 A probabilistic software quality model
abstract
In order to take the right decisions in estimating the costs and risks of a software change, it is crucial for the developers and managers to be aware of the quality attributes of their software. Maintainability is an important characteristic defined in the ISO/IEC 9126 standard, owing to its direct impact on development costs. Although the standard provides definitions for the quality characteristics, it does not define how they should be computed. Not being tangible notions, these characteristics are hardly expected to be representable by a single number. Existing quality models do not deal with ambiguity coming from subjective interpretations of characteristics, which depend on experience, knowledge, and even intuition of experts. This research aims at providing a probabilistic approach for computing high-level quality characteristics, which integrate expert knowledge, and deal with ambiguity at the same time. The presented method copes with “goodness” functions, which are continuous generalizations of threshold based approaches, i.e. instead of giving a number for the measure of goodness, it provides a continuous function. Two different systems were evaluated using this approach, and the results were compared to the opinions of experts involved in the development. The results show that the quality model values change in accordance with the maintenance activities, and they are in a good correlation with the experts' expectations.
Tibor Bakota, Péter Hegedüs, Peter Kortvelyesi, Rudolf Ferenc, Tibor Gyimóthy
ICSM5
2011 Adding Process Metrics to Enhance Modification Complexity Prediction
abstract
Software estimation is used in various contexts including cost, maintainability or defect prediction. To make the estimate, different models are usually applied based on attributes of the development process and the product itself. However, often only one type of attributes is used, like historical process data or product metrics, and rarely their combination is employed. In this report, we present a project in which we started to develop a framework for such complex measurement of software projects, which can be used to build combined models for different estimations related to software maintenance and comprehension. First, we performed an experiment to predict modification complexity (cost of a unity change) based on a combination of process and product metrics. We observed promising results that confirm the hypothesis that a combined model performs significantly better than any of the individual measurements.
Gabriella Tóth, Ádám Zoltán Végh, Árpád Beszédes, Tibor Gyimóthy
ICPC4
2010 Effect of test completeness and redundancy measurement on post release failures - An industrial experience report
abstract
In risk-based testing, compromises are often made to release a system in spite of knowing that it has outstanding defects. In an industrial setting, time and cost are often the “exit criteria” and - unfortunately - not the technical aspects like coverage or defect ratio. In such situations, the stakeholders accept that the remaining defects will be found after release, so sufficient resources are allocated to the “stabilization” phases following the release. It is hard for many organizations to see that such an approach is significantly costlier than trying to locate the defects earlier. We performed an empirical investigation of this for one of our industrial partners (a financial company). In this project, significant perfective maintenance was performed on the large information system. Based on changes made to the system, we carried out procedure level code coverage measurements with code level change impact analysis, and a similarity-based comparison of test cases in order to quantitatively check the completeness and redundancy of the tests performed. In addition, we logged and compared the number of defects found during testing and live operation. The data obtained were surprising for both the developers and the customer as well, leading to a major reorganization of their development, testing, and operation processes. After the reorganization, a significant improvement in these indicators for testing efficiency was observed.
Tamás Gergely, Árpád Beszédes, Tibor Gyimóthy, Milan Imre Gyalai
ICSM3
2010 MAGISTER: Quality assurance of Magic applications for software developers and end users
abstract
Nowadays there are many tools and methods available for source code quality assurance based on static analysis, but most of these tools focus on traditional software development techniques with 3GL languages. Besides procedural languages, 4GL programming languages such as Magic 4GL and Progress are widely used for application development. All these languages lie outside the main scope of analysis techniques. In this paper we present MAGISTER, which is a quality assurance framework for applications being developed in Magic, a 4GL application development solution created by Magic Software Enterprises. MAGISTER extracts data using static analysis methods from applications being developed in different versions of Magic (v5-9 and uniPaaS). The extracted data (including metrics, rule violations and dependency relations) is presented to the user via a GUI so it can be queried and visualized for further analysis. It helps software developers, architects and managers through the full development cycle by performing continuous code scans and measurements.
Csaba Nagy 0001, László Vidács, Rudolf Ferenc, Tibor Gyimóthy, Ferenc Kocsis
ICSM4
2010 New Conceptual Coupling and Cohesion Metrics for Object-Oriented Systems
abstract
The paper presents two novel conceptual metrics for measuring coupling and cohesion in software systems. Our first metric, Conceptual Coupling between Object classes (CCBO), is based on the well-known CBO coupling metric, while the other metric, Conceptual Lack of Cohesion on Methods (CLCOM5), is based on the LCOM5 cohesion metric. One advantage of the proposed conceptual metrics is that they can be computed in a simpler (and in many cases, programming language independent) way as compared to some of the structural metrics. We empirically studied CCBO and CLCOM5 for predicting fault-proneness of classes in a large open source system and compared these metrics with a host of existing structural and conceptual metrics for the same task. As the result, we found that the proposed conceptual metrics, when used in conjunction, can predict bugs nearly as precisely as the 58 structural metrics available in the Columbus source code quality framework and can be effectively combined with these metrics to improve bug prediction.
Bela Ujhazi, Rudolf Ferenc, Denys Poshyvanyk, Tibor Gyimóthy
SCAM4
2009 Modeling class cohesion as mixtures of latent topics
abstract
The paper proposes a new measure for the cohesion of classes in object-oriented software systems. It is based on the analysis of latent topics embedded in comments and identifiers in source code. The measure, named as maximal weighted entropy, utilizes the latent Dirichlet allocation technique and information entropy measures to quantitatively evaluate the cohesion of classes in software. This paper presents the principles and the technology that stand behind the proposed measure. Two case studies on a large open source software system are presented. They compare the new measure with an extensive set of existing metrics and use them to construct models that predict software faults. The case studies indicate that the novel measure captures different aspects of class cohesion compared to the existing cohesion measures and improves fault prediction for most metrics, which are combined with maximal weighted entropy.
Yixun Liu, Denys Poshyvanyk, Rudolf Ferenc, Tibor Gyimóthy, Nikos Chrisochoides
ICSM4
2009 Using information retrieval based coupling measures for impact analysis
Denys Poshyvanyk, Andrian Marcus, Rudolf Ferenc, Tibor Gyimóthy
Empir. Softw. Eng.4
2009 Combining preprocessor slicing with C/C++ language slicing
László Vidács, Árpád Beszédes, Tibor Gyimóthy
Sci. Comput. Program.3
2008 Static Execute After/Before as a replacement of traditional software dependencies
abstract
The paper explores Static Execute After (SEA) dependencies in the program and their dual Static Execute Before (SEB) dependencies. It empirically compares the SEA/SEB dependencies with the traditional dependencies that are computed by System Dependence Graph (SDG) and program slicers. In our case study we use about 30 subject programs that were previously used by other authors in empirical studies of program analysis. We report two main results. The computation of SEA/SEB is much less expensive and much more scalable than the computation of the SDG. At the same time, the precision declines only very slightly, by some 4% on average. In other words, the precision is comparable to that of the leading traditional algorithms, while intuitively a much larger difference would be expected. The paper then discusses whether based on these results the computation of the SDG should be replaced in some applications by the computation of the SEA/SEB.
Judit Jász, Árpád Beszédes, Tibor Gyimóthy, Václav Rajlich
ICSM3
2008 Combining Preprocessor Slicing with C/C++ Language Slicing
abstract
Slicing C programs has been one of the most popular ways for the implementation of slicing algorithms; out of the very few practical implementations that exist many deal with this programming language. Yet, preprocessor related issues have been addressed very marginally by these slicers, despite the fact that ignoring (or handling poorly) these constructs may lead to serious inaccuracies in the slicing results and hence in the comprehension process. Recently, an accurate slicing method for preprocessor related constructs has been proposed which - when combined with existing C/C++ language slicers - can provide a more complete comprehension of these languages. In this paper, we overview our approach for this combination and report its benefits in terms of the completeness of the resulting slices.
László Vidács, Judit Jász, Árpád Beszédes, Tibor Gyimóthy
ICPC4
2007 Clone Smells in Software Evolution
abstract
Although source code cloning (copy&paste programming) represents a significant threat to the maintainability of a software system, problems usually start to arise only when the system evolves. Most of the related research papers tackle the question of finding code clones in one particular version of the software only, leaving the dynamic behavior of the clones out of consideration. Eliminating these clones in large software systems often seems absolutely hopeless, as there might exist several thousands of them. Alternatively, tracking the evolution of individual clones can be used to identify those occurrences that could really cause problems in the future versions. In this paper we present an approach for mapping clones from one particular version of the software to another one, based on a similarity measure. This mapping is used to define conditions under which clones become suspicious (or "smelly") compared to their other occurrences. Accordingly, these conditions introduce the notion of dynamic clone smells. The usefulness of these smells is validated on the Mozilla Firefox internet browser, where the approach was able to find specific bugs that resulted from neglecting earlier copy&paste activities.
Tibor Bakota, Rudolf Ferenc, Tibor Gyimóthy
ICSM3
2007 Computation of Static Execute After Relation with Applications to Software Maintenance
abstract
In this paper, we introduce static execute after (SEA) relationship among program components and present an efficient analysis algorithm. Our case studies show that SEA may approximate static slicing with perfect recall and high precision, while being much less expensive and more usable. When differentiating between explicit and hidden dependencies, our case studies also show that SEA may correlate with direct and indirect class coupling. We speculate that SEA may find applications in computation of hidden dependencies and through it in many maintenance tasks, including change propagation and regression testing.
Árpád Beszédes, Tamás Gergely, Judit Jász, Gabriella Tóth, Tibor Gyimóthy, Václav Rajlich
ICSM5
2006 Towards Portable Metrics-based Models for Software Maintenance Problems
abstract
The usage of software metrics for various purposes has become a hot research topic in academia and industry (e.g. detecting design patterns and bad smells, studying change-proneness, quality and maintainability, predicting faults). Most of these topics have one thing in common: they are all using some kind of metrics-based models to achieve their goal. Unfortunately, only few researchers have tested these models on unknown software systems so far. This paper tackles the question, which metrics are suitable for preparing portable models (which can be efficiently applied to unknown software systems). We have assessed several metrics on four large software systems and we found that the well-known RFC and WMC metrics differentiate the analyzed systems fairly well. Consequently, these metrics cannot be used to build portable models, while the CBO, LCOM and LOC metrics behave similarly on all systems, so they seem to be suitable for this purpose
Tibor Bakota, Rudolf Ferenc, Tibor Gyimóthy, Claudio Riva, Jianli Xu
ICSM3
2006 Compacting XML documents
Miklós Kálmán, Ferenc Havasi, Tibor Gyimóthy
Inf. Softw. Technol.3
2006 A formalisation of the relationship between forms of program slicing
Dave W. Binkley, Sebastian Danicic, Tibor Gyimóthy, Mark Harman, Ákos Kiss 0001, Bogdan Korel
Sci. Comput. Program.3
2006 IEEE International Conference on Software Maintenance (ICSM2005)
Václav Rajlich, Tibor Gyimóthy
J. Softw. Maintenance Res. Pract.2
2006 Theoretical foundations of dynamic program slicing
Dave W. Binkley, Sebastian Danicic, Tibor Gyimóthy, Mark Harman, Ákos Kiss 0001, Bogdan Korel
Theor. Comput. Sci.3
2006 Guest Editors' Introduction to the Special Issue on the International Conference on Software Maintenance and Evolution
abstract
THE International Conference on Software Maintenance and Evolution (ICSM) is the leading research conference in software maintenance and evolution. It provides a widely recognized international forum for discussions and exchange of ideas between researchers and experts. ICSM 2005 was held in Budapest on 26-29 September 2005. In cooperation with the ICSM conference, several collocated workshops were organized. These included the Seventh IEEE International Symposium on Web Site Evolution (WSE), the Fifth IEEE International Workshop on Source Code Analysis and Manipulation (SCAM), the Third International Workshop on Visualizing Software for Understanding and Analysis (VISSOFT), the Software Evolvability Workshop, and the 13th International Workshop on Software Technology and Engineering Practice (STEP). For ICSM 2005, 180 research papers were submitted. From this record number of papers the program committee selected 55 full papers for inclusion in the proceedings. Of these 55 papers, eight were extended from their conference version, passed through the standard TSE review process involving three anonymous reviewers per paper, and are contained in this special issue. The first paper, “Feature Identification: An Epidemiological Metaphor” by G. Antoniol and Y.-G. Gueheneuc, presents an epidemiology inspired approach to feature identification. The approach relies on statistical analysis of static and dynamic data and identifies features of large, multithreaded object-oriented programs: the Firefox e-mail tool and the Mozilla Web browser. The results suggest that the epidemiological analysis overcomes the limitations of existing methods and increases the accuracy of the feature identification. Extracted microarchitectures consist of variables, structures, classes, and methods, and support the maintenance and program understanding. The second paper, “Toward the Reverse Engineering of UML Sequence Diagrams for Distributed Java Programs” by L.C. Briand, Y. Labiche, and J. Leduc, provides a methodology to reverse engineer sequence diagrams from dynamic analysis data. Because of inheritance, polymorphism, and dynamic binding, reverse engineering and understanding the behavior of an object-oriented system is difficult. Multithreading and distribution further complicate the analysis. In order to separate the instrumentation from the application code, aspect-oriented programming was used. The approach is illustrated on a distributed library management system developed in Java, using RMI as the distribution middleware. The third paper, “Static Analysis of Object References in RMI-Based Java Software” by M. Sharp and A. Rountev, presents a points-to analysis of distributed Java programs that contain the Remote Method Invocation (RMI). Points-to information is often a prerequisite for program understanding, testing, and code optimizations. The existing points-to analysis techniques cannot be applied directly to RMI-based distributed applications. The approach presented in this paper allows us to represent the flow of remote objects, the effects of remote invocations, and the remote propagation of object graphs through serialization. The result of an experimental study suggests that the approximate version of the analysis could be a good option for a relatively precise and practical points-to analysis of RMI-based Java applications. The fourth paper, “Incremental Maintenance of Software Artifacts” by S.P. Reiss, presents a software development tool (CLIME) that identifies potential problems and inconsistencies between software design, specification, documentation, source code, test cases, and other artifacts. The approach is based on constraints implemented as database queries that describe the association of the different software artifacts. These constraints are maintained and checked incrementally during the evolution of a system and the inconsistencies are reported to the developers. The approach is validated on several software development projects where many language and documentation problems and inconsistencies between the code and documentation have been identified. The fifth paper, “Tool-Supported Refactoring of ObjectOriented Code into Aspects” by D. Binkley, M. Ceccato, M. Harman, F. Ricca, and P. Tonella, presents refactoring Object-Oriented Programs (OOP), written in Java, into equivalent Aspect-Oriented Programs (AOP) written in AspectJ. The approach relies on six refactorings to support migration from OOP to AOP. These refactorings are combined with existing OO transformations into a tool, AOP-Migrator, which is implemented as an Eclipse plug-in. The approach is applied to several Java systems: JHotDraw, IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, VOL. 32, NO. 9, SEPTEMBER 2006 625
Tibor Gyimóthy, Václav Rajlich
IEEE Trans. Software Eng.1
2005 Using Dynamic Information in the Interprocedural Static Slicing of Binary Executables
Ákos Kiss 0001, Judit Jász, Tibor Gyimóthy
Softw. Qual. J.3
2005 Empirical Validation of Object-Oriented Metrics on Open Source Software for Fault Prediction
abstract
Open source software systems are becoming increasingly important these days. Many companies are investing in open source projects and lots of them are also using such software in their own work. But, because open source software is often developed with a different management style than the industrial ones, the quality and reliability of the code needs to be studied. Hence, the characteristics of the source code of these projects need to be measured to obtain more information about it. This paper describes how we calculated the object-oriented metrics given by Chidamber and Kemerer to illustrate how fault-proneness detection of the source code of the open source Web and e-mail suite called Mozilla can be carried out. We checked the values obtained against the number of bugs found in its bug database - called Bugzilla - using regression and machine learning methods to validate the usefulness of these metrics for fault-proneness prediction. We also compared the metrics of several versions of Mozilla to see how the predicted fault-proneness of the software system changed during its development cycle.
Tibor Gyimóthy, Rudolf Ferenc, István Siket
IEEE Trans. Software Eng.1
2004 Fact Extraction and Code Auditing with Columbus and SourceAudit
abstract
Automatic fact extraction from software systems is the fundamental building block in the process of understanding the relationships among a system's elements. We demonstrate the reverse engineering framework called Columbus which is able to automatically extract facts from C++ source code and how the extracted facts can be used in practice. We also mention a special-purpose tool that was developed on top of the Columbus framework. This tool, called SourceAudit, is a code auditor that is able to investigate source code and check it against rules that describe the preferred properties of the code.
Rudolf Ferenc, Árpád Beszédes, Tibor Gyimóthy
ICSM3
2004 Extracting Facts from Open Source Software
abstract
Open source software systems are becoming increasingly important these days. Many companies are investing in open source projects and lots of them are also using such software in their own work. But because open source software is often developed without proper management, the quality and reliability of the code may be uncertain. The quality of the code needs to be measured and this can be done only with the help of proper tools. We describe a framework called Columbus with which we calculate the object oriented metrics validated by Basili et al. for illustrating how fault-proneness detection from the open source Web and e-mail suite called Mozilla can be done. We also compare the metrics of several versions of Mozilla to see how the predicted fault-proneness of the software system changed during its development. The Columbus framework has been further developed recently with a compiler wrapping technology that now gives us the possibility of automatically analyzing and extracting information from software systems without modifying any of the source code or makefiles. We also introduce our fact extraction process here to show what logic drives the various tools of the Columbus framework and what steps need to be taken to obtain the desired facts.
Rudolf Ferenc, István Siket, Tibor Gyimóthy
ICSM3
2004 Seventh European Conference on Software Maintenance and Reengineering (CSMR 2003)
Mark van den Brand, Gerardo Canfora, Tibor Gyimóthy
J. Softw. Maintenance Res. Pract.3
2003 Annotated Hungarian National Corpus
Zoltán Alexin, János Csirik, Tibor Gyimóthy, Károly Bibok, Csaba Hatvani, Gábor Prószéky, László Tihanyi
EACL3
2002 Union Slices for Program Maintenance
abstract
Owing to their relative simplicity and wide range of applications, static slices are specifically proposed for software maintenance and program understanding. Unfortunately, in many cases static slices are overly conservative and therefore too large to supply useful information to the software maintainer. Dynamic slicing methods can produce more precise results, but only for one test case. In this paper we introduce the concept of union slices (the union of dynamic slices for many test cases) and suggest using a combination of static and union slices. This way the size of program parts that need to be investigated can be reduced by concentrating on the most important parts first. We performed a series of experiments with our experimental implementation on three medium size C programs. Our initial results suggest that union slices are in most cases far smaller than static slices, and that the growth rate of union slices (by adding more test cases) significantly declines after several representative executions of the program.
Árpád Beszédes, Csaba Faragó, Zsolt Mihály Szabó, János Csirik, Tibor Gyimóthy
ICSM5
2002 Columbus - Reverse Engineering Tool and Schema for C++
abstract
One of the most critical issues in large-scale software development and maintenance is the rapidly growing size and complexity of software systems. As a result of this rapid growth there is a need to better understand the relationships between the different parts of a large software system. In this paper we present a reverse engineering framework called Columbus that is able to analyze large C++ projects, and a schema for C++ that prescribes the form of the extracted data. The flexible architecture of the Columbus system with a powerful C++ analyzer and schema makes it a versatile and readily extendible toolset for reverse engineering. This tool is free for scientific and educational purposes and we fervently hope that it will assist academic persons in any research work related to C++ re- and reverse engineering.
Rudolf Ferenc, Árpád Beszédes, Mikko Tarkiainen, Tibor Gyimóthy
ICSM4
2002 Static and Dynamic Slicing of Constraint Logic Programs
Gyöngyi Szilágyi, Tibor Gyimóthy, Jan Maluszynski
Autom. Softw. Eng.2
1997 Application of Inductive Logic Programming for Learning ECG Waveforms
Gabriella Kókai, Zoltán Alexin, Tibor Gyimóthy
AIME3
1997 IMPUT: An Interactive Learning Tool Based on Program Specialization
abstract
The algorithm SPECTRE specializes logic programs with respect to positive and negative examples by applying the transformation rule unfolding together with clause removal. The method IMPUT presented in this paper gives a modified version of this algorithm by integrating the algorithmic debugging system IDTS with SPECTRE. The main idea of the IMPUT method, is that the identification of a clause to be unfolded has a crucial importance on the effectiveness of the specialization process. The debugging system IDTS is used to identify this buggy clause.
Zoltán Alexin, Tibor Gyimóthy, Henrik Boström
Intell. Data Anal.2
1996 Integrating Algorithmic Debugging and Unfolding Transformation in an Interactive Learner
Zoltán Alexin, Tibor Gyimóthy, Henrik Boström
ECAI2
1992 Integrated Graphics Environment to Develop Applications Based on Attribute Grammars
Tibor Gyimóthy, Zoltán Alexin, Róbert Szücs
CC1
1991 Generalized Algorithmic Debugging and Testing
abstract
This paper presents a method for semi-automatic bug localization, generalized algorithmic debugging, which has been integrated with the category partition method for functional testing. In this way the efficiency of the algorithmic debugging method for bug localization can be improved by using test specifications and test results. The long-range goal of this work is a semi-automatic debugging and testing system which can be used during large-scale program development of nontrivial programs. The method is generally applicable to procedural languages and is not dependent on any ad hoc assumptions regarding the subject program. The original form of algorithmic debugging, introduced by Shapiro, was however limited to small Prolog programs without side-effects, but has later been generalized to concurrent logic programming languages. Another drawback of the original method is the large number of interactions with the user during bug localization. To our knowledge, this is the first method which uses category partition testing to improve the bug localization properties of algorithmic debugging. The method can avoid irrelevant questions to the programmer by categorizing input parameters and then match these against test cases in the test database. Additionally, we use program slicing, a data flow analysis technique, to
Peter Fritzson, Tibor Gyimóthy, Mariam Kamkar, Nahid Shahmehri
PLDI2
1990 THALES: a Software Package for Plane Geometry Constructions with a Natural Language Interface
Károly Fábricz, Zoltán Alexin, Tibor Gyimóthy, Tamás Horváth 0001
COLING3