VLDB 2026 Research / reviewers in the wild / expert
Rainer Koschke
dblp:84/6899
· DBLP profile ↗
69ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0003-4094-3444ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 69 · 13 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluation of the Language Server Protocol for Static Dependency AnalysisabstractResearchers as well as practitioners often use static dependency graphs as a foundation for their investigations on the structure of a program. They are gathered by compiler-like static analyzers that are specific to a programming language. If more than one programming language is to be analyzed, different static analyzers need to be used, each having its own data structures and APIs, which increases the integration effort. To reduce this integration effort to a minimum, a standard mechanism to obtain dependency information would be of great help. The language server protocol (LSP) is such a standardized mechanism. It was developed in the context of multi-language integrated development environments (IDE) to implement interactive features such as auto-complete or code navigation. In this paper, we investigate whether LSP can be used to create static dependency graphs in a non-interactive way for$\mathrm{C}++, \mathrm{C} #$, Go, Java, JavaScript/TypeScript, Python, and Rust. Our use case differs from LSP's original purpose by the magnitude of the queries for nodes and edges, potentially bringing the language servers to their limits. We assess the scalability of the various language servers available for these languages and take a look at how various size metrics affect the different run-time phases. We found that LSP is a real help in integrating different tools to create static dependency graphs in a uniform way. Adding a different language server for a new language requires very little effort. Yet, scalability is a real issue. The gathering of the dependency data can take hours for large projects. The expected run-time for the analyses can be predicted as a linear function of the lines of code or number of files to be processed, allowing one to estimate in advance when a result can be expected. Falko Galperin, Michel Krause, Rainer Koschke |
ICSME | 3 |
| 2022 | Visualizing Code Smells: Tables or Code Cities? A Controlled ExperimentabstractThis paper presents a study in which we compared the visualization of code smells in Code Cities with classical tabular representations. We conducted a controlled experiment with 20 participants who had to solve six tasks in both environments. We evaluated the results of our experiment statistically and came to the following conclusions: In four tasks the completion time was significantly lower when using Code Cities (the remaining two tasks did not show any statistically significant differences for any of the two environments). Also, the perceived effort was significantly lower in three tasks when the participants used Code Cities over the tabular representation (again, the remaining tasks did not show any statistically significant differences). However, with regard to the perceived usability (which was measured across all tasks) and the correctness of the supplied answers, the tabular representation performed better. In particular, in two tasks the correctness was significantly better in the tabular environment and in one task the correctness was significantly better in the Code Cities environment. Based on our results, our verdict is as follows: Code Cities are better suited to get a quick overview of the code smells of a software, whereas tabular representations are better suited to analyze code smells in more detail.Experiment data, evaluation scripts and supplemental material: https://github.com/uni-bremen-agst/VISSOFT2022/archive/ refs/tags/1.0.0.zip (see README.md in the ZIP archive) Falko Galperin, Rainer Koschke, Marcel Steinbeck |
VISSOFT | 2 |
| 2022 | Edge Animation in Software VisualizationabstractAn important aspect in software visualization is the visualization of relations between elements. Examples include function calls across source-code files, software clones, and the like. A popular means of visualizing such relations are edges depicted as lines visually connecting the related elements. To diminish the visual clutter that can occur when drawing a large number of edges, hierarchical edge bundles, based on B-Splines, have proven suitable. In addition to a purely static view of software, animations, both of the elements and the edges, can contribute to the usability of a visualization by making changes easier to track for users. In this paper we discuss how B-Splines can be used as a basis for a variety of edge layouts and how edges represented as B-Splines can be animated using B-Spline morphing. We also discuss the challenges of morphing B-Splines, especially of those whose shape does not result solely from the positions of the connected elements (e.g., straight lines), but whose structure depends on multiple aspects (e.g., hierarchical edge bundles) and provide a solution to make arbitrary B-Splines compatible for morphing. The practicality of edge animation is demonstrated by use cases which we already implemented or plan to implement in the future in our software visualization tool SEE.Supplemental material: https://github.com/tinyspline/vissoft2022/archive/refs/tags/1.0.0.zip (see README.md in the zip archive) Marcel Steinbeck, Rainer Koschke |
VISSOFT | 2 |
| 2022 | The Effect of Feature Characteristics on the Performance of Feature Location TechniquesabstractFeature Location (FL)is a core software maintenance activity that aims to locate observable functionalities in the source code. Given its key role in software change, a vast array of Feature Location Techniques (FLTs) have been proposed but, as more and more FLTs are introduced, theselection of an appropriate FLTis an increasingly difficult problem. One consideration is thecharacteristics of the featuresbeing sought. For example, in the code associated with the feature, programmers may have named identifiers consistently, and with meaningful naming conventions, or not, and this may impact on the suitability of different FLTs. The suggestion that such characteristics matter has implicit support in the literature: An analysis of existing FLT empirical studies reveals that the system under study can often have a stronger impact on FLT performance than differing FLTs themselves. To understand this interaction between feature characteristics and FLTs better, this paper proposesa suite of feature-characteristic metricsthat are postulated to control FLTs’ performance, holistically across FLTs and impacting on individual FLTs to different degrees. To evaluate the suite, a controlled experiment is performed, using 878 features, to probe the relationship between the metrics and the performance of four FTL techniques: three commonly-used techniques and one state-of-the-art technique. The evaluation is performed using four commonly used evaluation measures and extended by employing 41 other established source-code metrics as extraneous variables. Results of the empirical evaluation suggest that the feature-metric suite presented impacts FLT performance holistically, and impacts different FLTs to different degrees. Thus, this paper moves towards the more standard selection of appropriate FLTs, with respect to the prominent feature characteristics in the software systems under study, and more rigorous consideration of the features selected to compare FLTs. Anthony Ventresque, Rainer Koschke, Andrea De Lucia, Jim Buckley |
IEEE Trans. Software Eng. | 3 |
| 2021 | Recording, Visualising and Understanding Developer Programming BehaviourabstractTo understand how developers solve programming tasks, it is necessary to observe what they are doing, i.e., what specific actions they perform, which strategies they apply and how they make use of possibly present, but not yet known or otherwise not yet available information that is needed to solve a task. To do so, we implemented a plug-in for the Eclipse IDE which captures nearly every interaction with the IDE a developer performs when working on a programming task. This enables us to comprehensively track a developer’s behaviour, e.g., whether and when a code edit needed for solving a given task was performed, and even more interesting, what were the preceding steps that led the developer to do so. In a first experiment conducted with the new plug-in, we were able to observe action patterns and program comprehension stages that confirm results of previous studies as well as were partly only suspected until now by recent literature, but never truly observed before.Screencast of the tool: https://youtu.be/GeZI-vCdgfoClosed caption version: https://youtu.be/mgt6Q-t7U00 Martin Schröer, Rainer Koschke |
SANER | 2 |
| 2021 | Javadoc Violations and Their Evolution in Open-Source SoftwareabstractSoftware quality comprises different and interrelated aspects. One of them is maintainability, which in turn is made up of measurable attributes. Previous studies have shown that documentation, by contributing to the comprehensibility of software, may have a positive effect on maintainability and, hence, software quality. This paper presents a study in which we analyzed Javadoc comments from 163 different open-source projects. Javadoc is the de facto standard for documenting source code files in Java projects, and although its syntax is less strict than in other (programming) languages, documentation written with Javadoc may contain violations. Our study focuses on the detection of different types of Javadoc violations as well as the source code elements affected by them. Also, by utilizing software repository mining techniques, we examined the history of the subject systems to gain further insights into the evolution of Javadoc violations. According to our results, about half of the source code elements have no Javadoc whatsoever. Among the different components of Javadoc comments (if present), the description of exceptions, by far, has the highest average ratio of violations. With regard to the types of affected elements, constructors and methods show very high average ratios. Also, we found that, on average, violations live more than two years.Nowadays, most integrated development environments (IDEs) for Java are capable of detecting missing Javadoc comments as well as comments with syntactic errors. However, our results indicate that the documentation of source code might be considered less important to developers or that these tools alone may not be sufficient for maintaining consistent documentation. Marcel Steinbeck, Rainer Koschke |
SANER | 2 |
| 2021 | TinySpline: A Small, yet Powerful Library for Interpolating, Transforming, and Querying NURBS, B-Splines, and Bézier CurvesabstractNURBS, B-Splines, and Bézier curves have a wide range of applications. One of them is software visualization, where these kinds of splines are often used to depict relations between objects in a visually appealing manner. For example, several visualization techniques make use of hierarchical edge bundles to reduce the visual clutter that can occur when drawing a large number of edges. Another example is the visualization of software in 3D, virtual reality, and augmented reality environments. In these environments edges can be drawn as splines in 3D space to overcome the natural limitations of the two-dimensional plane-e.g., the collision of edges with other objects. While Bézier curves are supported quite well by most UI frameworks and game engines, NURBS and B-Splines are not. Hence, spline-based visualizations are considerably more difficult to implement without in-depth knowledge in the area of splines.In this paper we present TinySpline, a general purpose library for NURBS, B-Splines, and Bézier curves that is well suited for implementing advanced edge visualization techniques-e.g., but not limited to, hierarchical edge bundles. The core of the library is written in ANSI C with a C++ wrapper for an object-oriented programming model. Based on the C++ wrapper, auto-generated bindings for C#, D, Go, Java, Lua, Octave, PHP, Python, R, and Ruby are provided, which enables TinySpline to be integrated into a large number of applications. Marcel Steinbeck, Rainer Koschke |
SANER | 2 |
| 2020 | Static Extraction of Enforced Authorization Policies SeeAuthzabstractAuthorization is an intrinsic part of a software's security. Determining whether a user is allowed to access a resource or not is crucial, not only in safety-critical applications but also in everyday applications to prevent misuse of data or software. There is plenty of research dealing with validating and verifying authorization policies in the security community. Still, an implemented authorization policy does not necessarily match the planned authorization policy, i.e., even a validated and verified authorization policy can pose security issues when implemented incorrectly. This gap between planned and implemented authorization policy poses the risk of unauthorized access to sensitive resources due to insufficient authorization checks. Therefore, it is essential to ensure a system's security to validate the implemented authorization policy against the planned one. We, therefore, describe the authorization pattern and present an algorithm to extract authorization graphs from implemented authorization policies, which can then be used to compare against the planned authorization policy. To that end, we developed a configurable context-sensitive analysis tailored to Java-based software systems, where the context is the authorization facts that hold on each point. Using a configuration for Apache Shiro, a security library that supports authorization, we evaluated our implementation using an open-source repository system for the management and dissemination of digital content and a closed-source manufacturing execution system. We discuss additional usage scenarios of the analysis results and describe how to transfer the approach to other authorization policies and programming languages. Bernhard J. Berger, Rodrigue Wete Nguempnang, Karsten Sohr, Rainer Koschke |
SCAM | 4 |
| 2020 | Clustering Paths With Dynamic Time WarpingabstractStudying software visualization often includes the evaluation of paths collected from participants of a study (e.g., eye tracking or movements in virtual worlds). In this paper, we explore clustering techniques to automate the process of grouping similar paths. The heart of the evaluated approach is a distance metric between paths that is based on dynamic time warping (DTW). DTW aligns two paths based on any given distance metric between their data points so as to minimize the distance between those paths-alignment may stretch or compress time for best fit. With a data set of 127 paths of professional software developers exploring code cities in virtual reality, we evaluate the clustering based on objective quality indices and manual inspection. Rainer Koschke, Marcel Steinbeck |
VISSOFT | 1 |
| 2020 | How EvoStreets Are Observed in Three-Dimensional and Virtual Reality EnvironmentsabstractWhen analyzing software systems, a large amount of data accumulates. In order to assist developers in the preparation, evaluation, and understanding of findings, different visualization techniques have been developed. Due to recent progress in immersive virtual reality, existing visualization tools were ported to this environment. However, three-dimensional and virtual reality environments have different advantages and disadvantages, and by transferring concepts, such as layout algorithms and user interaction mechanisms, more or less one-to-one, the characteristics of these environments are neglected. In order to develop techniques adapting to the circumstance of a particular environment, more research in this field is necessary. In previously conducted case studies, we compared EvoStreets deployed in three different environments: 2D, 2.5D, and virtual reality. We found evidence that movement patterns—path length, average speed, and occupied volume—differ significantly between the 2.5D and virtual reality environments for some of the tasks that had to be solved by 34 participants in a controlled experiment. In this paper, we analyze the results of this experiment in more details, to study if not only movement is affected by these environments, but also the way how EvoStreets are observed. Although we could not find enough evidence that the number of viewpoints and their duration differ significantly, we found indications that in virtual reality viewpoints are located closer to the EvoStreets and that the distance between viewpoints is shorter. Based on our previous results and the findings of this paper, we present visualization and user interaction concepts specific to the kind of environment. Marcel Steinbeck, Rainer Koschke, Marc O. Rüdel |
SANER | 2 |
| 2020 | Mining understandable state machine models from embedded code
Wasim Said, Jochen Quante, Rainer Koschke |
Empir. Softw. Eng. | 3 |
| 2019 | Do extracted state machine models help to understand embedded software?abstractProgram understanding is a prerequisite for several software activities, such as maintenance, evolution, and reengineering. Code in itself is so detailed that it is often hard to understand. More abstract models describing its behaviour may ease program understanding. Manually building understandable abstractions from complex source code - as an explicit or just mental model - requires in-depth analysis of the code in the first place. Therefore, it is a time-consuming and tedious activity for developers. Model mining can support program comprehension by semi-automatically extracting high-level models from code. One helpful model is a state machine, which is an established formalism for specifying the behaviour of a software component. In this paper, we report on a controlled experiment that investigates the question: Do semi-automatically extracted state machines make understanding of complex embedded code more effective? The experiment was conducted with 30 participants on two industrial embedded C code functions. The results show that the share of correct answers increases and the required time to solve the tasks decreases significantly when extracted state machines are available. We conclude that mined state machines do in fact help in program understanding. Wasim Said, Jochen Quante, Rainer Koschke |
ICPC | 3 |
| 2019 | Comparing the EvoStreets visualization technique in two D and three-dimensional environments: a controlled experimentabstractAnalyzing and maintaining large software systems is a challenging task due to the sheer amount of information contained therein. To overcome this problem, Steinbrückner developed a visualization technique named EvoStreets. Utilizing the city metaphor, EvoStreets are well suited to visualize the hierarchical structure of a software as well as hotspots regarding certain aspects. Early implementations of this approach use three-dimensional rendering on regular two-dimensional displays. Recently, though, researchers have begun to visualize EvoStreets in virtual reality using head-mounted displays, claiming that this environment enhances user experience. Yet, there is little research on comparing the differences of EvoStreets visualized in virtual reality with EvoStreets visualized in conventional environments. This paper presents a controlled experiment, involving 34 participants, in which we compared the EvoStreet visualization technique in different environments, namely, orthographic projection with keyboard and mouse, 2.5D projection with keyboard and mouse, and virtual reality with head-mounted displays and hand-held controllers. Using these environments, the participants had to analyze existing Java systems regarding software clones. According to our results, it cannot be assumed that: 1) the orthographic environment takes less time to find an answer, 2) the 2.5D and virtual reality environments provide better results regarding the correctness of edge-related tasks compared to the orthographic environment, and 3) the performance regarding time and correctness differs between the 2.5D and virtual reality environments. Marcel Steinbeck, Rainer Koschke, Marc O. Rüdel |
ICPC | 2 |
| 2019 | The Architectural Security Tool Suite - ARCHSECabstractArchitectural risk analysis is a risk management process for identifying security flaws at the level of software architectures and is used by large software vendors, to secure their products. We present our architectural security environ- ment (ARCHSEC) that has been developed at our institute during the past eight years in several research projects. ARCHSEC aims to simplify architectural risk analysis, making it easier for small and mid-sized companies to get started. With ARCHSEC, it is possible to graphically model or to reverse engineer software security architectures. The regained software architectures can then be inspected manually or au- tomatically analyzed w.r.t. security flaws, resulting in a threat model, which serves as a base for discussion between software and security experts to improve the overall security of the software system in question, beyond the level of implementation bugs. In the evaluation part of this paper, we demonstrate how we use ARCHSEC in two of our current research projects to analyze business applications. In the first project we use ARCHSEC to identify security flaws in business process diagrams. In the second project, ARCHSEC is integrated into an audit environment for software security certification. ARCHSEC is used to identify security flaws and to visualize software systems to improve the effectiveness and efficiency of the certification process. Bernhard J. Berger, Karsten Sohr, Rainer Koschke |
SCAM | 3 |
| 2019 | Movement Patterns and Trajectories in Three-Dimensional Software VisualizationabstractSoftware visualization is a growing field of research, in which developers are assisted in understanding and analyzing complex applications by mapping different aspects of a software system onto visual attributes. Under the assumption that virtual reality, due to the higher degree of immersion, may enhance user experience, researchers have begun to port existing visualization techniques to this environment. Oftentimes, layout algorithms and user interaction methods are more or less transferred one-to-one, though little is known about the effect of virtual reality in visual analytics and program comprehension. Moreover, little research on the behavior of developers in different three-dimensional visualization environments has been done yet. This paper extends the results of a previous controlled experiment, in which the EvoStreets visualization technique was compared in different two-and three-dimensional environments. In the original experiment, we could not find evidence that any of the environments, namely, 2D, 2.5D, and virtual reality, effects the time required to find an answer or the correctness of the given answer. However, we found indications that movement patterns differ between the 2.5D and the virtual reality environments. For this paper, we analyzed and refined the movement trajectories that have been recorded in the previous experiment. We found significant differences for some of the tasks that had to be solved by the participants. In particular, we found evidence that the path length, average speed, and occupied volume differ. Though we could find significant correlations between these metrics and correctness, we found indications that there is a correlation with time, which, in turn, differs significantly between the 2.5D and the VR environments for many tasks. These findings may have implications on the design of visualizations, interactions, and recommendation systems for these different environments. Marcel Steinbeck, Rainer Koschke, Marc O. Rüdel |
SCAM | 2 |
| 2019 | Towards Understandable Guards of Extracted State Machines from Embedded SoftwareabstractThe extraction of state machines from complex software systems can be very useful to understand the behavior of a software, which is a prerequisite for other software activities, such as maintenance, evolution and reengineering. However, using static analysis to extract state machines from real-world embedded software often leads to models that cannot be understood by humans: The extracted models contain a high number of states and transitions and very complex guards (transition conditions). Integrating user interaction into the extraction process can reduce these state machines to an acceptable size. However, the problem of highly complex guards remains. In this paper, we present a novel approach to reduce the complexity of guards in such state machines to a degree that is understandable for humans. The conditions are reduced by a combination of heuristic logic minimization, masking of infeasible paths, and using transition priorities. The approach is evaluated with software developers on industrial embedded C code. The results show that the approach is highly effective in making the guards understandable. Our controlled experiment shows that guards reduced by our approach and presented with priorities are easier to understand than guards without priorities. Wasim Said, Jochen Quante, Rainer Koschke |
SANER | 3 |
| 2018 | On State Machine Mining from Embedded Control SoftwareabstractProgram understanding is a time-consuming and tedious activity for software developers. Manually building abstractions from source code requires in-depth analysis of the code in the first place. Model mining can support program comprehension by semi-automatically extracting high-level models from code. One potentially helpful model is a state machine, which is an established formalism for specifying the behavior of a software component. There exist only few approaches for state machine mining, and they either deal with object-oriented systems or expect specific state implementation patterns. Both preconditions are usually not met by real-world embedded control software written in procedural languages. Other approaches extract only API protocols instead of the component's behavior. In this paper, we propose and evaluate several techniques that enable state machine mining from embedded control code: 1) We define criteria for state variables in procedural code based on an empirical study. This enables adaptation of an existing approach for extracting state machines from object-oriented software to embedded control code. 2) We present a refinement of the transition extraction process of that approach by removing infeasible transitions, which on average leads to more than 50% reduction of the number of transitions. 3) We evaluate two approaches to reduce the complexity of transition conditions. 4) An empirical study examines the limits of transition conditions' complexity that can still be understood by humans. These techniques and studies constitute major building blocks towards mining understandable state machines from embedded control software. Wasim Said, Jochen Quante, Rainer Koschke |
ICSME | 3 |
| 2018 | Reflexion Models for State Machine Extraction and VerificationabstractHigh-level design models are often used for describing the behavior or structure of a software system. It is generally much easier and more adequate to understand a software system on this level than on the level of individual code lines. Such models are also created by developers as they gain an understanding of the software. Unfortunately, these models often do not correspond to what is really in the code. Murphy et al. introduced the idea of reflexion models in 1995 to overcome this problem. Their approach is today widely used for architecture conformance checking and reconstruction. In this paper, we introduce reflexion models for state machines. Our approach allows to check the correspondence of a hypothetical state machine model with the code. It returns information about convergence, partial convergence, divergence, or absence of the specified states and transitions. Similar to the original reflexion model, the approach can be used for conformance checking as well as interactive reverse engineering of state machine models. We concentrate on the latter and show the potential of the approach in several case studies. Wasim Said, Jochen Quante, Rainer Koschke |
ICSME | 3 |
| 2018 | Towards Interactive Mining of Understandable State Machine Models from Embedded Software
Wasim Said, Jochen Quante, Rainer Koschke |
MODELSWARD | 3 |
| 2018 | [Engineering Paper] Built-in Clone Detection in Meta LanguagesabstractDevelopers often practice re-use by copying and pasting code. Copied and pasted code is also known as clones. Clones may be found in all programming languages. Automated clone detection may help to detect clones in order to support software maintenance and language design. Syntax-based clone detectors find similar syntax subtrees and, hence, are guaranteed to yield only syntactic clones. They are also known to have high precision and good recall. Developing a syntax-based clone detector for each language from scratch may be an expensive task. In this paper, we explore the idea to integrate syntax-based clone detection into workbenches for language engineering. Such workbenches allow developers to create their own domain-specific language or to create parsers for existing languages. With the integration of clone detection into these workbenches, a clone detector comes as a free byproduct of the grammar specification. The effort is spent only once for the workbench and not multiple times for every language built with the workbench. We report our lessons learned in applying this idea for three language workbenches: the popular parser generator ANTLR and two language workbenches for domain-specific languages, namely, MPS, developed by JetBrains, and Xtext, which is based on the Eclipse Modeling Framework. Rainer Koschke, Urs-Bjorn Schmidt, Bernhard J. Berger |
SCAM | 1 |
| 2018 | A Controlled Experiment on Spatial Orientation in VR-Based Software CitiesabstractMultiple authors have proposed a city metaphor for visualizing software. While early approaches have used three-dimensional rendering on standard two-dimensional displays, recently researchers have started to use head-mounted displays to visualize software cities in immersive virtual reality systems (IVRS). For IVRS of a higher order it is claimed that they offer a higher degree of engagement and immersion as well as more intuitive interaction. On the other hand, spatial orientation may be a challenge in IVRS as already reported by studies on the use of IVRS in domains outside of software engineering such as gaming, education, training, and mechanical engineering or maintenance tasks. This might be even more true for the city metaphor for visualizing software. Software is immaterial and, hence, has no natural appearance. Only a limited number of abstract aspects of software are mapped onto visual representations so that software cities generally lack the details of the real world, such as the rich variety of objects or fine textures, which are often used as clues for orientation in the real world. In this paper, we report on an experiment in which we compare navigation in a particular kind of software city (EvoStreets) in two variants of IVRS. One with head-mounted display and hand controllers versus a 3D desktop visualization on a standard display with keyboard and mouse interaction involving 20 participants. Marc O. Rüdel, Johannes Ganser, Rainer Koschke |
VISSOFT | 3 |
| 2016 | Special section on software clones
Nils Göde, Yoshiki Higo, Rainer Koschke |
Softw. Qual. J. | 3 |
| 2015 | From preprocessor-constrained parse graphs to preprocessor-constrained control flowabstractPreprocessor-aware static analysis tools are needed for C Code to gain sound knowledge about the interference among all conditionally compiled program parts. We provide formal descriptions and algorithms to construct a preprocessor-aware control flow graph from preprocessor-aware parse graphs of SuperC. Based on the structure of parse graphs capturing the syntax nodes constrained by preprocessor constraints, we show how to model, formalize, and compute preprocessor-aware intra-procedural control-flow graphs. Such preprocessor-aware control-flow graphs may serve as the basis for subsequent preprocessor-aware control and data flow analyses. Dierk Lüdemann, Rainer Koschke |
SCAM | 2 |
| 2015 | A survey on goal-oriented visualization of clone dataabstractComprehending software clones is necessary for a number of activities in software development. The comprehension of software clones is challenged by the sheer volume of data and the complexity of the information content in that data. Visualization, or visual data analysis, takes advantage of human cognitive skills to discover unstructured insights from the visual presentations of complex and voluminous data. In this paper, we survey the existing literature on visualization of software clones. We gather the insights provided, and put that information in context of actual information needs systematically derived from the clone management goals. This framework allows us to better understand the role a visualization may play in achieving a specific user goal, identify potential gaps between existing types of visualization and information needs, and find complementary non-redundant subsets of visualizations for each user goal. Hamid Abdul Basit, Muhammad Hammad 0001, Rainer Koschke |
VISSOFT | 3 |
| 2015 | Preface special section on software clones (IWSC'13)
Rainer Koschke, Juergen Rilling |
J. Softw. Evol. Process. | 1 |
| 2014 | Effect of Clone Information on the Performance of Developers Fixing Cloned BugsabstractDuplicated source code -- clones -- is known to occur frequently in software systems and bears the risk of inconsistent updates of the code. The impact of clones has been investigated mostly by retrospective analysis of software systems. Only little effort has been spent to investigate human interaction when dealing with clones. A previous study by Chatterji and colleagues found that cloned defects are removed significantly more accurately when clone information is provided to the programmers. We conducted a controlled experiment to extend the previous study on the use of clone information by investigating the effect of clone information on the performance of developers in common bug-fixing tasks. The experiment shows that developers are quite capable to compensate missing clone information through testing to provide correct solutions. Clone information does help to detect cloned defects faster, although developers may exploit semantic code relations such as inheritance to uncover cloned defects only slightly slower if they do not have clone information. If cloned defects lurk in semantically unrelated places however, clone information helps to find them faster at statistical significance. Developers without clone information needed 17 minutes longer on average or 140% more time in relative terms to complete the task successfully. Saman Bazrafshan, Rainer Koschke |
SCAM | 2 |
| 2014 | Special issue on software clones (IWSC'12)
Katsuro Inoue, Rainer Koschke, Jens Krinke |
Sci. Comput. Program. | 2 |
| 2014 | Large-scale inter-system clone detection using suffix trees and hashingabstractSUMMARY Detecting a similar code between two systems has various applications such as comparing two software variants or versions or finding potential license violations. Techniques detecting suspiciously similar code must scale in terms of resources needed to very large code corpora and need to have high precision because a human needs to inspect the results. This paper demonstrates how suffix trees can be used to obtain a scalable comparison. The evaluation is carried out for very large code corpora. Our evaluation shows that our approach is faster than index‐based techniques when the analysis is run only once. If the analysis is to be conducted multiple times, creating an index pays off. We report how much code can be filtered out from the analysis using an index‐based filter. In addition to that, this paper proposes a method to improve precision through user feedback. A user validates a sample of the found clone candidates. An automated data mining technique learns a decision tree on the basis of the user decisions and different code metrics. We investigate the relevance of several metrics and whether criteria learned from one application domain can be generalized to other domains. Copyright © 2013 John Wiley & Sons, Ltd. Rainer Koschke |
J. Softw. Evol. Process. | 1 |
| 2014 | On the Comprehension of Program ComprehensionabstractResearch in program comprehension has evolved considerably over the past decades. However, only little is known about how developers practice program comprehension in their daily work. This article reports on qualitative and quantitative research to comprehend the strategies, tools, and knowledge used for program comprehension. We observed 28 professional developers, focusing on their comprehension behavior, strategies followed, and tools used. In an online survey with 1,477 respondents, we analyzed the importance of certain types of knowledge for comprehension and where developers typically access and share this knowledge. We found that developers follow pragmatic comprehension strategies depending on context. They try to avoid comprehension whenever possible and often put themselves in the role of users by inspecting graphical interfaces. Participants confirmed that standards, experience, and personal communication facilitate comprehension. The team size, its distribution, and open-source experience influence their knowledge sharing and access behavior. While face-to-face communication is preferred for accessing knowledge, knowledge is frequently shared in informal comments. Our results reveal a gap between research and practice, as we did not observe any use of comprehension tools and developers seem to be unaware of them. Overall, our findings call for reconsidering the research agendas towards context-aware tool support. Walid Maalej, Rebecca Tiarks, Tobias Roehm, Rainer Koschke |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2013 | 7th international workshop on software clones (IWSC 2013)abstractSoftware Clones are identical or similar pieces of code, models or designs. In this, the 7th International Workshop on Software Clones (IWSC), we will discuss issues in software clone detection, analysis and management, as well as applications to software engineering contexts that can benefit from knowledge of clones. These are important emerging topics in software engineering research and practice. Special emphasis will be given this time to clone management in practice, emphasizing use cases and experiences. We will also discuss broader topics on software clones, such as clone detection methods, clone classification, management, and evolution, the role of clones in software system architecture, quality and evolution, clones in plagiarism, licensing, and copyright, and other topics related to similarity in software systems. The format of this workshop will give enough time for intense discussions. Rainer Koschke, Elmar Jürgens, Juergen Rilling |
ICSE | 1 |
| 2013 | An Empirical Study of Clone RemovalsabstractIt is often claimed that duplicated source code is a threat to the maintainability of a software system and that developers should manage code duplication. A previous study analyzed the evolution of four software systems and found a remarkable discrepancy between code clones detected by a state-of-the-art clone detector and those deliberately removed by developers as the scope of the clones hardly ever matched. However, the results are based on a relatively small amount of data and need to be validated by a more extensive analysis. In this paper, we present an extension of this study by analyzing deliberate as well as accidental removals of code duplication in the evolution of eleven systems. Based on our findings, we could confirm the results of the previous study. Beyond that we found that accidental removals of cloned code occur slightly more often than deliberate removals and that many clone removals were in fact incomplete. Saman Bazrafshan, Rainer Koschke |
ICSM | 2 |
| 2013 | Studying clone evolution using incremental clone detectionabstractSUMMARY Finding, understanding and managing software clones—passages of duplicated source code—is of large interest in research and practice. Analyzing the evolution of clones across multiple versions of a program adds value to both applications. Although there is an abundance of techniques to detect clones, current approaches are limited to a single version of a program. The current techniques to track clones utilize these single‐version approaches and map clones of consecutive versions retroactively. This causes an unnecessary overhead in runtime and may lead to an incorrect mapping due to ambiguity. In this paper, we present an incremental clone detection algorithm, which detects clones based on the results of the previous version's analysis. It creates a mapping between clones of consecutive versions along with the detection. We evaluated our incremental approach regarding its advantage in runtime as well as the usefulness of the mapping for studies on the clone evolution. Copyright © 2010 John Wiley & Sons, Ltd. Nils Göde, Rainer Koschke |
J. Softw. Evol. Process. | 2 |
| 2013 | Incremental reflexion analysisabstractSUMMARY Architecture conformance checking is implemented in many commercial and research tools. These tools typically implement the reflexion analysis originally proposed by Murphy, Notkin, and Sullivan. This analysis allows for structural validation of an architecture model against a source model connected by a mapping from source entities onto architecture entities. Given this mapping, the reflexion analysis computes the discrepancies between the architecture model and source model automatically. The mapping process is usually highly interactive and the most time‐consuming activity in the reflexion analysis. In current tools, the reflexion analysis must be repeated completely whenever the underlying source or architecture models or the mapping changes. In large systems, the recomputation can hinder interactive use as users expect an immediate response to their changes. This paper describes an incremental reflexion analysis that does not require a complete repetition of the reflexion analysis. Instead, it repeats the analysis only for those parts that are actually influenced by a change. The incremental reflexion analysis is evaluated on large real‐world systems. Copyright © 2011 John Wiley & Sons, Ltd. Rainer Koschke |
J. Softw. Evol. Process. | 1 |
| 2012 | How do professional developers comprehend software?abstractResearch in program comprehension has considerably evolved over the past two decades. However, only little is known about how developers practice program comprehension under time and project pressure, and which methods and tools proposed by researchers are used in industry. This paper reports on an observational study of 28 professional developers from seven companies, investigating how developers comprehend software. In particular we focus on the strategies followed, information needed, and tools used. We found that developers put themselves in the role of end users by inspecting user interfaces. They try to avoid program comprehension, and employ recurring, structured comprehension strategies depending on work context. Further, we found that standards and experience facilitate comprehension. Program comprehension was considered a subtask of other maintenance tasks rather than a task by itself. We also found that face-to-face communication is preferred to documentation. Overall, our results show a gap between program comprehension research and practice as we did not observe any use of state of the art comprehension tools and developers seem to be unaware of them. Our findings call for further careful analysis and for reconsidering research agendas. Tobias Roehm, Rebecca Tiarks, Rainer Koschke, Walid Maalej |
ICSE | 3 |
| 2012 | Program complexity metrics and programmer opinionsabstractVarious program complexity measures have been proposed to assess maintainability. Only relatively few empirical studies have been conducted to back up these assessments through empirical evidence. Researchers have mostly conducted controlled experiments or correlated metrics with indirect maintainability indicators such as defects or change frequency. This paper uses a different approach. We investigate whether metrics agree with complexity as perceived by programmers. We show that, first, programmers' opinions are quite similar and, second, only few metrics and in only few cases reproduce complexity rankings similar to human raters. Data-flow metrics seem to better match the viewpoint of programmers than control-flow metrics, but even they are only loosely correlated. Moreover we show that a foolish metric has similar or sometimes even better correlation than other evaluated metrics, which raises the question how meaningful the other metrics really are. In addition to these results, we introduce an approach and associated statistical measures for such multi-rater investigations. Our approach can be used as a model for similar studies. Bernhard Katzmarski, Rainer Koschke |
ICPC | 2 |
| 2011 | Fifth international workshop on software clones: (IWSC 2011)abstractSoftware clones are identical or similar pieces of code, design or other artifacts. Clones are known to be closely related to various issues in software engineering, such as software quality, complexity, architecture, refactoring, evolution, licensing, plagiarism, and so on. Various characteristics of software systems can be uncovered through clone analysis, and system restructuring can be performed by merging clones. James R. Cordy, Katsuro Inoue, Stan Jarzabek, Rainer Koschke |
ICSE | 4 |
| 2011 | Frequency and risks of changes to clonesabstractCode Clones - duplicated source fragments - are said to increase maintenance effort and to facilitate problems caused by inconsistent changes to identical parts. While this is certainly true for some clones and certainly not true for others, it is unclear how many clones are real threats to the system's quality and need to be taken care of. Our analysis of clone evolution in mature software projects shows that most clones are rarely changed and the number of unintentional inconsistent changes to clones is small. We thus have to carefully select the clones to be managed to avoid unnecessary effort managing clones with no risk potential. Nils Göde, Rainer Koschke |
ICSE | 2 |
| 2011 | Guest editor's introduction to the special section on the 2009 international conference on program comprehension (ICPC 2009)
Rainer Koschke, Andrian Marcus, Gerald C. Gannod |
Softw. Qual. J. | 1 |
| 2011 | An extended assessment of type-3 clones as detected by state-of-the-art tools
Rebecca Tiarks, Rainer Koschke, Raimar Falke |
Softw. Qual. J. | 2 |
| 2010 | Fourth International Workshop on Software Clones (IWSC)abstractSoftware clones are identical or similar pieces of code. They are often the result of copy--and--paste activities as ad-hoc code reuse by programmers. Software clones research is of high relevance for the industry. Many researchers have reported high rates of code cloning in both industrial and open-source systems. Katsuro Inoue, Stan Jarzabek, James R. Cordy, Rainer Koschke |
ICSE (2) | 4 |
| 2009 | An Assessment of Type-3 Clones as Detected by State-of-the-Art ToolsabstractCode reuse through copying and pasting leads to so-called software clones. These clones can be roughly categorized into identical fragments (type-1 clones), fragments with parameter substitution (type-2 clones), and similar fragments that differ through modified,deleted, or added statements (type-3 clones). Although there has been extensive research on detecting clones, detection of type-3 clones is still an open research issue due to the inherent vaguenessin their definition. In this paper, we analyze type-3 clones detected by state-of-the-art tools and investigate type-3 clones in terms of their syntactic differences. Then, we derive their underlying semantic abstractions from their syntactic differences. Finally, we investigate whether there are any additional code characteristics that indicate that a tool-suggested clone candidate is a real type-3 clone from a human's perspective. Our findings can help developers of clone detectors to improve their tools. Rebecca Tiarks, Rainer Koschke, Raimar Falke |
SCAM | 2 |
| 2009 | Comparison and evaluation of code clone detection techniques and tools: A qualitative approach
Chanchal Kumar Roy, James R. Cordy, Rainer Koschke |
Sci. Comput. Program. | 3 |
| 2009 | An evaluation of code similarity identification for the grow-and-prune modelabstractAbstract In case new functionality is required, which is similar to the existing one, developers often copy the code that implements the existing functionality and adjust the copy to the new requirements. The result of the copying is code growth. If developers face maintenance problems, because of the need to make changes multiple times for the original and all its copies, they may decide to merge the original and its copies again; that is, they prune the code. This approach was named the grow‐and‐prune model by Faust and Verhoef. This paper describes tool support for the grow‐and‐prune model in the evolution of software by identifying similar functions that may be merged. These functions are identified in two steps. First, token‐based clone detection is used to detect pairs of functions sharing code. Second, Levenshtein distance (LD) measures the textual similarity among these functions. Sufficient similarity at function level is then lifted to the architectural level. The approach is evaluated by a case study for the Linux kernel. We give examples of instances of the grow‐and‐prune model for Linux. Then, we evaluate our technique quantitatively by measuring recall and precision with respect to an oracle. To obtain the oracle, we asked nine different developers to decide whether they believe certain functions are similar and should be merged. The evaluation shows that the recall and precision of our technique are about 75%. Calculating LD on token values rather than characters is superior. The two metrics strongly correlate but the token‐based calculation reduces runtime by a factor of 4.6. Clone detection is an effective filter to reduce the number of calculations of the relatively expensive LD. Copyright © 2009 John Wiley & Sons, Ltd. Thilo Mende, Rainer Koschke, Felix Beckwermert |
J. Softw. Maintenance Res. Pract. | 2 |
| 2009 | Extending the reflexion method for consolidating software variants into product lines
Rainer Koschke, Pierre Frenzel, Andreas P. J. Breu, Karsten Angstmann |
Softw. Qual. J. | 1 |
| 2009 | A Systematic Survey of Program Comprehension through Dynamic AnalysisabstractProgram comprehension is an important activity in software maintenance, as software must be sufficiently understood before it can be properly modified. The study of a program's execution, known as dynamic analysis, has become a common technique in this respect and has received substantial attention from the research community, particularly over the last decade. These efforts have resulted in a large research body of which currently there exists no comprehensive overview. This paper reports on a systematic literature survey aimed at the identification and structuring of research on program comprehension through dynamic analysis. From a research body consisting of 4,795 articles published in 14 relevant venues between July 1999 and June 2008 and the references therein, we have systematically selected 176 articles and characterized them in terms of four main facets: activity, target, method, and evaluation. The resulting overview offers insight in what constitutes the main contributions of the field, supports the task of identifying gaps and opportunities, and has motivated our discussion of several important research directions that merit additional consideration in the near future. Bas Cornelissen, Andy Zaidman, Arie van Deursen, Leon Moonen, Rainer Koschke |
IEEE Trans. Software Eng. | 5 |
| 2008 | Empirical evaluation of clone detection using syntax suffix trees
Raimar Falke, Pierre Frenzel, Rainer Koschke |
Empir. Softw. Eng. | 3 |
| 2008 | Dynamic object process graphs
Jochen Quante, Rainer Koschke |
J. Syst. Softw. | 2 |
| 2008 | Encapsulating targeted component abstractions using software Reflexion ModellingabstractAbstract Design abstractions such as components, modules, subsystems or packages are often not made explicit in the implementation of legacy systems. Indeed, often the abstractions that are made explicit turn out to be inappropriate for future evolution agendas. This can make the maintenance, evolution and refactoring of these systems difficult. In this publication, we carry out a fine‐grained evaluation of Reflexion Modelling as a technique for encapsulating user‐targeted components. This process is a prelude to component recovery, reuse and refactoring. The evaluation takes the form of twoin vivocase studies, where two professional software developers encapsulate components in a large, commercial software system. The studies demonstrate the validity of this approach and offer several best‐use guidelines. Specifically, they argue that users benefit from having a strong mental model of the system in advance of Reflexion Modelling, even if that model is flawed, and that users should expend effort exploring the expected relationships present in Reflexion Models. Copyright © 2008 John Wiley & Sons, Ltd. Jim Buckley, Andrew Le Gear, Christopher Exton, Ross Cadogan, Trevor Johnston, Bill Looby, Rainer Koschke |
J. Softw. Maintenance Res. Pract. | 7 |
| 2007 | Automated clustering to support the reflexion method
Andreas Christl, Rainer Koschke, Margaret-Anne D. Storey |
Inf. Softw. Technol. | 2 |
| 2007 | Comparison and Evaluation of Clone Detection ToolsabstractMany techniques for detecting duplicated source code (software clones) have been proposed in the past. However, it is not yet clear how these techniques compare in terms of recall and precision as well as space and time requirements. This paper presents an experiment that evaluates six clone detectors based on eight large C and Java programs (altogether almost 850 KLOC). Their clone candidates were evaluated by one of the authors as independent third party. The selected techniques cover the whole spectrum of the state-of-the-art in clone detection. The techniques work on text, lexical and syntactic information, software metrics, and program dependency graphs. Stefan Bellon, Rainer Koschke, Giuliano Antoniol, Jens Krinke, Ettore Merlo |
IEEE Trans. Software Eng. | 2 |
| 2007 | Guest Editors' Introduction to the Special Section from the International Conference on Software Maintenance and EvolutionabstractSOFTWARE maintenance and evolution are relevant to users, engineers, and researchers who come into contact with software beyond Version 1. The International Conference on Software Maintenance and Evolution (ICSM) is the premiere forum for software maintenance researchers and practitioners to examine, discuss, and exchange ideas regarding the key issues facing the software maintenance community. During the conference, participants from academia, government, and industry share ideas and experiences solving critical software maintenance problems. ICSM 2006 was held in Philadelphia on 24-27 September 2006 in cooperation with several colocated workshops. These included the Eighth IEEE International Symposium on Web Site Evolution (WSE), the Sixth IEEE International Workshop on Source Code Analysis and Manipulation (SCAM), the Second International IEEE Workshop on Software Evolvability, and the Second International Workshop on Predictive Models of Modern Industrial Software Engineering. ICSM 2006’s technical program was anchored by 45 papers selected from 147 submissions. The program also included keynote addresses from three distinguished speakers: Patrick Lardieri, David Notkin, and Richard Stallman. Of the 45 papers, seven were invited to this special issue with five passing the rigorous TSE review process. These five selected papers include two that discuss frameworks and three that present new techniques. The frameworks support the creation of language-independent program analyses and evolvable test suites. Two of the techniques consider source-code reengineering and the final one the challenging problem of making live updates to a software system while it is running. The first of the two framework papers, “An Extensible Metamodel for Program Analysis” by D. Strein, R. Lincke, J. Lundberg, and W. Lowe, describes a language-independent framework for building new analyses. The resulting architecture supports the construction of specific analyses (e.g., refactorings) and is easily extensible through the addition of new front ends to support new languages. This flexibility and the use of loose coupling between components of the existing framework support the easy integration of new components. The paper uses the implementation of the tools VIZZANALYZER and XDEVELOP as a proof of concept. Looking forward, further work includes empirical study (e.g., considering running time and memory consumption) and the incorporation of dynamic analysis into the framework (e.g., to support debuggers and profilers). This paper was published in the September 2007 issue of TSE; however, the abstract of this paper is included in this special section. The paper can be found in the Computer Society Digital Library at http://computer.org/tse/ archives.htm. In the second framework paper, “On the Detection of Test Smells: A Metrics-Based Approach for General Fixture and Eager Test” by B. Van Rompaey, B. Du Bois, S. Demeyer, and M. Rieger, a framework for evaluating the evolvability of white box tests suites alongside the program is presented. The goal of this approach is to avoid the (often significant) cost incurred when test cases must be coevolved with the program. The framework accomplishes this by describing a collection of test smells. Akin to code smells, test smells allow the concrete expression of what a good evolvable test is by exploiting the principles that underlie white box testing. The paper proposes a formal description of test smells by means of metric predictors, empirically evaluated for two example test smells. Looking forward, future work will consider the interplay between test smells and frequently changing test cases. Understanding how test smells emerge and grow should help in proactively detecting current and future smells. The third and fourth papers address reengineering, a problem into which significant software evolution energy is invested. The first of these two papers, “API-Evolution Support with Diff-CatchUp” by Z. Xing and E. Stroulia, addresses the API evolution problem. In short, reusable components, and thus their APIs, must evolve in response to client needs for improved functionality, quality, and generality. Applications using an API which evolve independently can thus “break.” To tackle the API-evolution problem, the paper presents a technique and a tool that automatically recognizes API changes in a reused component and proposes plausible fixes based on working examples of the framework code base. Looking forward, planned extensions to the work IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, VOL. 33, NO. 12, DECEMBER 2007 797 Dave W. Binkley, Rainer Koschke, Spiros Mancoridis |
IEEE Trans. Software Eng. | 2 |
| 2006 | Introduction
Rick Kazman, Arie van Deursen, Rainer Koschke |
Autom. Softw. Eng. | 3 |
| 2006 | Selected papers from the fourth Source Code Analysis and Manipulation (SCAM 2004) Workshop
Tom Dean, Mark Harman, Rainer Koschke, Michael L. Van de Vanter |
J. Syst. Softw. | 3 |
| 2006 | Revisiting the Delta IC approach to component recovery
Rainer Koschke, Gerardo Canfora, Jörg Czeranski |
Sci. Comput. Program. | 1 |
| 2005 | On dynamic feature locationabstractFeature location aims at locating pieces of code that implement a given set of features (requirements). It is a necessary first step in every program comprehension and maintenance task if the connection between features and code has been lost.We have developed a semi-automatic technique for feature location using a combination of static and dynamic program analysis. Formal concept analysis is used to explore the results of the dynamic analysis.We describe new experiences with our technique. Specifically, we investigate the gain of information and increase of costs when the system under analysis is profiled at basic block level rather than routine level as in our earlier work. Furthermore, we explore the influence of the scenarios used for the dynamic analysis (minimal versus combined scenarios). Rainer Koschke, Jochen Quante |
ASE | 1 |
| 2005 | What Architects Should Know About Reverse Engineering and RengineeringabstractArchitecture reconstruction is a form of reverse engineering that reconstructs architectural views from an existing system. It is often necessary because a complete and authentic architectural description is not available. This paper puts forward the goals of architecture reconstruction, revisits the technical difficulties we are facing in architecture reconstruction, and presents a summary of a literature survey about the types of architectural viewpoints addressed in reverse engineering research. Rainer Koschke |
WICSA | 1 |
| 2005 | Static object trace extraction for programs with pointers
Thomas Eisenbarth 0003, Rainer Koschke, Gunther Vogel |
J. Syst. Softw. | 2 |
| 2004 | Symphony: View-Driven Software Architecture ReconstructionabstractAuthentic descriptions of a software architecture are required as a reliable foundation for any but trivial changes to a system. Far too often, architecture descriptions of existing systems are out of sync with the implementation. If they are, they must be reconstructed. There are many existing techniques for reconstructing individual architecture views, but no information about how to select views for reconstruction, or about process aspects of architecture reconstruction in general. In this paper we describe view-driven process for reconstructing software architecture that fills this gap. To describe Symphony, we present and compare different case studies, thus serving a secondary goal of sharing real-life reconstruction experience. The Symphony process incorporates the state of the practice, where reconstruction is problem-driven and uses a rich set of architecture views. Symphony provides a common framework for reporting reconstruction experiences and for comparing reconstruction approaches. Finally, it is a vehicle for exposing and demarcating research problems in software architecture reconstruction. Arie van Deursen, Christine Hofmeister, Rainer Koschke, Leon Moonen, Claudio Riva |
WICSA | 3 |
| 2004 | Addendum to "Locating Features in Source Code'abstractFor original paper by T. Eisenbarth et al. see ibid., vol.29, no.3, p.210-24 (2003). We compare three approaches that apply formal concept analysis on execution profiles. This survey extends the discourse of related research by Bojic and Velasevic (2000). Dragan Bojic, Thomas Eisenbarth 0003, Rainer Koschke, Daniel Simon, Dusan M. Velasevic |
IEEE Trans. Software Eng. | 3 |
| 2003 | Software visualization in software maintenance, reverse engineering, and re-engineering: a research surveyabstractAbstract Software visualization is concerned with the static visualization as well as the animation of software artifacts, such as source code, executable programs, and the data they manipulate, and their attributes, such as size, complexity, or dependencies. Software visualization techniques are widely used in the areas of software maintenance, reverse engineering, and re‐engineering, where typically large amounts of complex data need to be understood and a high degree of interaction between software engineers and automatic analyses is required. This paper reports the results of a survey on the perspectives of 82 researchers in software maintenance, reverse engineering, and re‐engineering on software visualization. It describes to which degree the researchers are involved in software visualization themselves, what is visualized and how, whether animation is frequently used, whether the researchers believe animation is useful at all, which automatic graph layouts are used if at all, whether the layout algorithms have deficiencies, and—last but not least—where the medium‐term and long‐term research in software visualization should be directed. The results of this survey help to ascertain the current role of software visualization in software engineering from the perspective of researchers in these domains and give hints on future research avenues. Copyright © 2003 John Wiley & Sons, Ltd. Rainer Koschke |
J. Softw. Maintenance Res. Pract. | 1 |
| 2003 | Locating Features in Source CodeabstractUnderstanding the implementation of a certain feature of a system requires identification of the computational units of the system that contribute to this feature. In many cases, the mapping of features to the source code is poorly documented. In this paper, we present a semiautomatic technique that reconstructs the mapping for features that are triggered by the user and exhibit an observable behavior. The mapping is in general not injective; that is, a computational unit may contribute to several features. Our technique allows for the distinction between general and specific computational units with respect to a given set of features. For a set of features, it also identifies jointly and distinctly required computational units. The presented technique combines dynamic and static analyses to rapidly focus on the system's parts that relate to a specific set of features. Dynamic information is gathered based on a set of scenarios invoking the features. Rather than assuming a one-to-one correspondence between features and scenarios as in earlier work, we can now handle scenarios that invoke many features. Furthermore, we show how our method allows incremental exploration of features while preserving the "mental map" the analyst has gained through the analysis. Thomas Eisenbarth 0003, Rainer Koschke, Daniel Simon |
IEEE Trans. Software Eng. | 2 |
| 2002 | Incremental Location of Combined Features for Large-Scale ProgramsabstractThe need for changing a program frequently confronts maintainers with the reality that no valid architectural description is at hand. To solve that problem, we presented at ICSM 2001 a language-independent and easy to use technique for opportunistic and demand driven location of features in source code based on static and dynamic analysis and concept analysis. In order to further validate the technique, we performed an industrial case study on a 1.2 million LOC production system. The experiences we made during that case study showed two problems of our approach: the growing complexity of concept lattices for large systems with many features and the need for handling compositions of features. This paper extends our technique to solve these problems. We show how this method allows incremental exploration of features while preserving the "mental map" the maintainer has gained through the analysis. The second improvement is a detailed look at composing features into more complex scenarios. Rather than assuming a one-to-one correspondence between features and scenarios as in earlier work, we can now handle scenarios that invoke many features. Thomas Eisenbarth 0003, Rainer Koschke, Daniel Simon |
ICSM | 2 |
| 2002 | Atomic Architectural Component Recovery for Program Understanding and EvolutionabstractComponent recovery and remodularization is a means to get back control on large and complex legacy systems suffering from ad-hoc changes by recovering logical components and restructuring the physical components accordingly to decrease coupling among components and increase cohesion of components. This thesis is on unifying, quantitatively and qualitatively evaluating, improving, and integrating automatic and semi-automatic methods for component recovery. Rainer Koschke |
ICSM | 1 |
| 2001 | Aiding Program Comprehension by Static and Dynamic Feature AnalysisabstractUnderstanding a system's implementation without prior knowledge is a hard task for reengineers in general. However, some degree of automatic aid is possible. The authors present a technique for building a mapping between the system's externally visible behavior and the relevant parts of the source code. The technique combines dynamic and static analyses to rapidly focus on the system's parts urgently required for a goal-directed process of program understanding. Thomas Eisenbarth 0003, Rainer Koschke, Daniel Simon |
ICSM | 2 |
| 2000 | Workshop on standard exchange format (WoSEF)abstractNo abstract available. Susan Elliott Sim, Richard C. Holt, Rainer Koschke |
ICSE | 3 |
| 2000 | A comparison of abstract data types and objects recovery techniques
Jean-Francois Girard, Rainer Koschke |
Sci. Comput. Program. | 2 |
| 1999 | A Metric-Based Approach to Detect Abstract Data Types and State Encapsulations
Jean-Francois Girard, Rainer Koschke, Georg Schied |
Autom. Softw. Eng. | 2 |
| 1997 | Finding Components in a Hierarchy of Modules: a Step towards Architectural UnderstandingabstractThis paper presents a method to view a system as a hierarchy of modules according to information hiding concepts and to identify architectural component candidates in this hierarchy. The result of the method eases the understanding of a system's underlying software architecture. A prototype tool implementing this method was applied to three systems written in C (each over 30 Kloc). For one of these systems, an author of the system created an architectural description. The components generated by our method correspond to those of this architectural description in almost all cases. For the other two systems, most of the components resulting from the method correspond to meaningful system abstractions Jean-Francois Girard, Rainer Koschke |
ICSM | 2 |
| 1997 | A Metric-based Approach to Detect Abstract Data Types and State EncapsulationsabstractThis article presents an approach to identify abstract data types (ADT) and abstract state encapsulations (ASE, also called abstract objects) in source code. This approach groups together functions, types, and variables into ADT and ASE candidates according to the proportion of features they share. The set of features considered includes the context of these elements, the relationships to their environment, and informal information. A prototype tool has been implemented to support this approach. It has been applied to three C systems (each between 30-38 Kloc). The ADTs and ASEs identified by the approach are compared to those identified by software engineers who did not know the proposed approach. In a case study, this approach has been shown to identify, in most cases, more ADTs and ASEs than five published techniques applied on the same systems. This is important when trying to identify as many ADTs and ASEs as possible. Jean-Francois Girard, Rainer Koschke, Georg Schied |
ASE | 2 |