Jurgen J. Vinju

dblp:55/5625 · DBLP profile ↗
← Back
49ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-2686-7409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 48 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2025 OIL: an industrial case study in language engineering with Spoofax
abstract
Abstract Domain-specific languages (DSLs) promise to improve the software engineering process, e.g., by reducing software development and maintenance effort and by improving communication, and are therefore seeing increased use in industry. To support the creation and deployment of DSLs, language workbenches have been developed. However, little is published about the actual added value of a language workbench in an industrial setting, compared to not using a language workbench. In this paper, we evaluate the productivity of using the Spoofax language workbench by comparing two implementations of an industrial DSL, one in Spoofax and one in Python, that already existed before the evaluation. The subject is the Open Interaction Language (OIL): a complex DSL for implementing control software with requirements imposed by its industrial context at Canon Production Printing. Our findings indicate that it is more productive to implement OIL using Spoofax compared to using Python, especially if editor services are desired. Although Spoofax was sufficient to implement OIL, we find that Spoofax should especially improve on practical aspects to increase its adoptability in industry.
Olav Bunte, Jasper Denkers, Louis C. M. van Gool, Jurgen J. Vinju, Eelco Visser, Tim A. C. Willemse, Andy Zaidman
Softw. Syst. Model.4
2023 Taming complexity of industrial printing systems using a constraint-based DSL: An industrial experience report
abstract
Abstract Flexible printing systems are highly complex systems that consist of printers, that print individual sheets of paper, and finishing equipment, that processes sheets after printing, for example, assembling a book. Integrating finishing equipment with printers involves the development of control software that configures the devices, taking hardware constraints into account. This control software is highly complex to realize due to (1) the intertwined nature of printing and finishing, (2) the large variety of print products and production options for a given product, and (3) the large range of finishers produced by different vendors. We have developed a domain‐specific language called CSX that offers an interface to constraint solving specific to the printing domain. We use it to model printing and finishing devices and to automatically derive constraint solver‐based environments for automatic configuration. We evaluate CSX on its coverage of the printing domain in an industrial context, and we report on lessons learned on using a constraint‐based DSL in an industrial context.
Jasper Denkers, Marvin Brunner, Louis C. M. van Gool, Jurgen J. Vinju, Andy Zaidman, Eelco Visser
Softw. Pract. Exp.4
2022 Breaking bad? Semantic versioning and impact of breaking changes in Maven Central
Lina Ochoa, Thomas Degueule, Jean-Rémy Falleri, Jurgen J. Vinju
Empir. Softw. Eng.4
2022 Large-scale semi-automated migration of legacy C/C++ test code
abstract
Abstract This is an industrial experience report on a large semi‐automated migration of legacy test code in C and C++. The particular migration was enabled by automating most of the maintenance steps. Without automation this particular large‐scale migration would not have been conducted, due to the risks involved in manual maintenance (risk of introducing errors, risk of unexpected rework, and loss of productivity). We describe and evaluate the method of automation we used on this real‐world case. The benefits were that by automating analysis, we could make sure that we understand all the relevant details for the envisioned maintenance, without having to manually read and check our theories. Furthermore, by automating transformations we could reiterate and improve over complex and large scale source code updates, until they were “just right.” The drawbacks were that, first, we have had to learn new metaprogramming skills. Second, our automation scripts are not readily reusable for other contexts; they were necessarily developed for this ad‐hoc maintenance task. Our analysis shows that automated software maintenance as compared to the (hypothetical) manual alternative method seems to be better both in terms of avoiding mistakes and avoiding rework because of such mistakes. It seems that necessary and beneficial source code maintenance need not to be avoided, if software engineers are enabled to create bespoke (and ad‐hoc) analysis and transformation tools to support it.
Mathijs Schuts, Rodin Aarssen, Paul Tielemans, Jurgen J. Vinju
Softw. Pract. Exp.4
2021 Modeling with Mocking
abstract
Writing formal specifications often requires users to abstract from the original problem. Especially when verification techniques such as model checking are used. Without applying abstraction the search space the model checker need to traverse tends to grow quickly beyond the scope of what can be checked within reasonable time.The downside of this need to omit details is that it increases the distance to the implementation. Ideally, the created specifications could be used to generate software from (either manually or automatically). But having an incomplete description of the desired system is not enough for this purpose.In this work we introduce the Rebel2 specification language. Rebel2 lets the user write full system specifications in the form of state machines with data without the need to apply abstraction while still preserving the ability to verify non-trivial properties. This is done by allowing the user to forget and mock specifications when running the model checker. The original specifications are untouched by these techniques.We compare the expressiveness of Rebel2 and the effectiveness of mock and forget by implementing two case studies: one from the automotive domain and one from the banking domain. We find that Rebel2 is expressive enough to implement both case studies in a concise manner. Next to that, when performing checks in isolation, mocking can speed up model checking significantly.
Jouke Stoel, Tijs van der Storm, Jurgen J. Vinju
ICST3
2021 Getting grammars into shape for block-based editors
abstract
Block-based environments are visual programming environments that allow users to program by interactively arranging visual jigsaw-like blocks. They have shown to be helpful in several domains but often require experienced developers for their creation. Previous research investigated the use of language workbenches to generate block-based editors based on grammars, but the generated block-based editors sometimes provided too many unnecessary blocks, leading to verbose environments and programs. To reduce the number of interactions, we propose a set of transformations to simplify the original grammar, yielding a reduction of the number of (useful) kinds of blocks available in the resulting editors. We show that our generated block-based editors are improved for a set of observed aesthetic criteria up to a certain complexity. As such, analyzing and simplifying grammars before generating block-based editors allows us to derive more compact and potentially more usable block-based editors, making reuse of existing grammars through automatic generation feasible.
Mauricio Verano Merino, Tom Beckmann, Tijs van der Storm, Robert Hirschfeld, Jurgen J. Vinju
SLE5
2019 Rascal, 10 Years Later
abstract
The following topics are dealt with: software maintenance; public domain software; program debugging; program diagnostics; program testing; software quality; Java; text analysis; program compilers; software engineering.
Paul Klint, Tijs van der Storm, Jurgen J. Vinju
SCAM3
2018 An empirical evaluation of OSGi dependencies best practices in the eclipse IDE
abstract
OSGi is a module system and service framework that aims to fill Java's lack of support for modular development. Using OSGi, developers divide software into multiple bundles that declare constrained dependencies towards other bundles. However, there are various ways of declaring and managing such dependencies, and it can be confusing for developers to choose one over another. Over the course of time, experts and practitioners have defined "best practices" related to dependency management in OSGi. The underlying assumptions are that these best practices (i) are indeed relevant and (ii) help to keep OSGi systems manageable and efficient. In this paper, we investigate these assumptions by first conducting a systematic review of the best practices related to dependency management issued by the OSGi Alliance and OSGi-endorsed organizations. Using a large corpus of OSGi bundles (1,124 core plug-ins of the Eclipse IDE), we then analyze the use and impact of 6 selected best practices. Our results show that the selected best practices are not widely followed in practice. Besides, we observe that following them strictly reduces classpath size of individual bundles by up to 23% and results in up to ±13% impact on performance at bundle resolution time. In summary, this paper contributes an initial empirical validation of industry-standard OSGi best practices. Our results should influence practitioners especially, by providing evidence of the impact of these best practices in real-world systems.
Lina Ochoa, Thomas Degueule, Jurgen J. Vinju
MSR3
2018 To-many or to-one? all-in-one! efficient purely functional multi-maps with type-heterogeneous hash-tries
abstract
An immutable multi-map is a many-to-many map data structure with expected fast insert and lookup operations. This data structure is used for applications processing graphs or many-to-many relations as applied in compilers, runtimes of programming languages, or in static analysis of object-oriented systems. Collection data structures are assumed to carefully balance execution time of operations with memory consumption characteristics and need to scale gracefully from a few elements to multiple gigabytes at least. When processing larger in-memory data sets the overhead of the data structure encoding itself becomes a memory usage bottleneck, dominating the overall performance.
Michael J. Steindorfer, Jurgen J. Vinju
PLDI2
2018 Bacatá: a language parametric notebook generator (tool demo)
abstract
Interactive notebooks allow people to communicate and collaborate through a single rich document that might include live code, multimedia, computed results, and documentation, which is persisted as a whole for reproducibility. Notebooks are currently being used extensively in domains such as data science, data journalism, and machine learning. However, constructing a notebook interface for a new language requires a lot of effort. In this tool paper, we present Bacatá, a language parametric notebook generator for domain-specific languages (DSL) based on the Jupyter framework. Bacatá is designed so that language engineers may reuse existing language components (such as parsers, code generators, interpreters, etc.) as much as possible. Moreover, we explain the design of Bacatá and how DSL notebooks can be generated with minimum effort in the context of the Rascal meta programming system and language workbench.
Mauricio Verano Merino, Jurgen J. Vinju, Tijs van der Storm
SLE2
2017 Challenges for static analysis of Java reflection: literature review and empirical study
abstract
The behavior of software that uses the Java Reflection API is fundamentally hard to predict by analyzing code. Only recent static analysis approaches can resolve reflection under unsound yet pragmatic assumptions. We survey what approaches exist and what their limitations are. We then analyze how real-world Java code uses the Reflection API, and how many Java projects contain code challenging state-of-the-art static analysis. Using a systematic literature review we collected and categorized all known methods of statically approximating reflective Java code. Next to this we constructed a representative corpus of Java systems and collected descriptive statistics of the usage of the Reflection API. We then applied an analysis on the abstract syntax trees of all source code to count code idioms which go beyond the limitation boundaries of static analysis approaches. The resulting data answers the research questions. The corpus, the tool and the results are openly available. We conclude that the need for unsound assumptions to resolve reflection is widely supported. In our corpus, reflection can not be ignored for 78% of the projects. Common challenges for analysis tools such as non-exceptional exceptions, programmatic filtering meta objects, semantics of collections, and dynamic proxies, widely occur in the corpus. For Java software engineers prioritizing on robustness, we list tactics to obtain more easy to analyze reflection code, and for static analysis tool builders we provide a list of opportunities to have significant impact on real Java code.
Davy Landman, Alexander Serebrenik, Jurgen J. Vinju
ICSE3
2017 Guest editors' introduction to the 6th issue of Experimental Software and Toolkits (EST-6)
Mark van den Brand, Jurgen J. Vinju, Kim Mens
Sci. Comput. Program.2
2017 Enabling PHP software engineering research in Rascal
Mark Hills 0001, Paul Klint, Jurgen J. Vinju
Sci. Comput. Program.3
2017 Corrigendum: Empirical analysis of the relationship between CC and SLOC in a large corpus of Java methods and C functions published on 9 December 2015
abstract
INTRODUCTION During the preparation of the corresponding chapter in Davy Landman's PhD thesis, some minor graphical and statistical discrepancies were found in the paper “Empirical analysis of the relationship between CC and SLOC in a large corpus of Java methods and C functions.” To support future reproduction and use of this work, we prepared the current erratum, containing several updated figures, a diagnosis of the cause of the errors, and an explanation of the effect on the original paper. None of the issues reported in this erratum influence the conclusions of the original paper. ISSUES DISCOVERED The hexagonal scatter plots in Figure lack a more prominent line at CC = 0. This was caused by a bug *reported and confirmed: https://github.com/tidyverse/ggplot2/issues/2061in ggplot, which would filter out data around the limits. The R2 values in the Tables B and B of the C corpus were off by a maximum of 0.01 from the actual result. The cause was that this table was not re-calculated after fixing a bug in the “remove out-of-scope code” phase. Note that the impact of this error is scattered throughout the paper, as the correlations of Tables and are often repeated for clarity in the remaining sections (for example, the R2 of the linear model for all the C functions is 0.43 instead of 0.44). Our R code calculating the log-transformed linear fit contained an error. The dashed lines in Figures, and are impacted and the shape of the residual plot in 11. The biggest impact is in Figure, where the original fit seemed to miss the data almost entirely. We misinterpreted this phenomenon in the last sentence of the second paragraph of section 4.4.2; it is not caused by the skewness of the distributions of the two metrics, but rather by the current bug. The custom implementation of the log-scaled y-axis of the residual plots in Figure contained two errors: ∘ The labels on the y-axis were off by a factor 10 ∘For the negative side of the residual plot, we took the absolute, calculated the log10 value, and made it negative again. However, values between 0 and 1 (the values close to the linear fit) turn into a negative value (as log10(1) equals 0). This caused strange outliers in the original plots that were not scrutinized. The fixed residual plots do not have this outliers and look much more like the data in Figure. We republished the data sets related to the current paper on Zenodo to increase their availability: ∘ Landman, Davy. (2015). A Curated Corpus of Java Source Code based on Sourcerer (2015) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.208213 ∘Landman, Davy. (2015). A Large Corpus of C Source Code based on Gentoo packages [Data set]. Zenodo. http://doi.org/10.5281/zenodo.208215 ∘Davy Landman. (2015, February 26). cwi-swat/jsep-sloc-versus-cc. Zenodo. http://doi.org/10.5281/zenodo.293795 NEW IMAGES The remaining part of this erratum contains updated tables and figures as replacements for the original paper. 8 (Figure presented.) Scatter plots of SLOC vs CC zoomed in on the bottom left quadrant. The solid and dashed lines are the linear regression before and after the log transform. The grayscale gradient of the hexagons is logarithmic 4 Correlations for part of the tail of the independent variable SLOC. All correlations have a high significance level p≤1×10−16.(b) C functions (Table presented.) 5 Correlations for part of the tail of the independent variable SLOC removed. All correlations have a high significance level p≤1×10−16.(b) C functions (Table presented.) 9 (Figure presented.) Scatter plots of SLOC vs CC on a log-log scale. The solid and dashed lines are the linear regression before and after the log transform. The grayscale gradient of the hexagons is logarithmic 11 (Figure presented.) Residual plot of the linear regressions after the log transform, both axis are on a log scale. The grayscale gradient of the hexagons is logarithmic 12 (Figure presented.) Scatter plots of SLOC vs CC for Java and C files. The solid and dashed lines are the linear regression before and after the log transform. The grayscale gradient of the hexagons is logarithmic.
Davy Landman, Alexander Serebrenik, Eric Bouwers, Jurgen J. Vinju
J. Softw. Evol. Process.4
2016 Towards a software product line of trie-based collections
abstract
Collection data structures in standard libraries of programming languages are designed to excel for the average case by carefully balancing memory footprint and runtime performance. These implicit design decisions and hard-coded trade-offs do constrain users from using an optimal variant for a given problem. Although a wide range of specialized collections is available for the Java Virtual Machine (JVM), they introduce yet another dependency and complicate user adoption by requiring specific Application Program Interfaces (APIs) incompatible with the standard library.
Michael J. Steindorfer, Jurgen J. Vinju
GPCE2
2016 Towards a universal code formatter through machine learning
Terence Parr, Jurgen J. Vinju
SLE2
2016 On Error-Class Distribution in Automotive Model-Based Software
abstract
Software fault prediction promises to be a powerful tool in supporting test engineers upon their decision where to define testing hotspots. However, there are limitations on a cross project prediction and a lack of reports upon application to industrial software, as well as the power of metrics to represent bugs. In this paper, we present a novel analysis based upon faults discovered in model-based automotive software projects and their relationship to metrics used to perform fault prediction. Using our previously released dataset on software metrics, we report bug classes discovered during heavy testing of those automotive software. As the software has been developed following strict coding and development guidelines, we present the results based on a comparison between the discovered error classes and those which might derive a reduced potential error set. Using the three projects from our dataset we determine if any of these bug classes are project specific.
Harald Altinger, Yanjindulam Dajsuren, Sebastian Siegl, Jurgen J. Vinju, Franz Wotawa
SANER4
2016 Performance Modeling of Maximal Sharing
abstract
It is noticeably hard to predict the effect of optimization strategies in Java without implementing them. "Maximal sharing" (a.k.a. "hash-consing") is one of these strategies that may have great benefit in terms of time and space, or may have detrimental overhead. It all depends on the redundancy of data and the use of equality. We used a combination of new techniques to predict the impact of maximal sharing on existing code: Object Redundancy Profiling (ORP) to model the effect on memory when sharing all immutable objects, and Equals-Call Profiling (ECP) to reason about how removing redundancy impacts runtime performance. With comparatively low effort, using the MAximal SHaring Oracle (MASHO), a prototype profiler based on ORP and ECP, we can uncover optimization opportunities that otherwise would remain hidden.
Michael J. Steindorfer, Jurgen J. Vinju
ICPE2
2016 Empirical analysis of the relationship between CC and SLOC in a large corpus of Java methods and C functions
abstract
Abstract Measuring the internal quality of source code is one of the traditional goals of making software development into an engineering discipline. Cyclomatic complexity (CC) is an often used source code quality metric, next to source lines of code (SLOC). However, the use of the CC metric is challenged by the repeated claim that CC is redundant with respect to SLOC because of strong linear correlation. We conducted an extensive literature study of the CC/SLOC correlation results. Next, we tested correlation on large Java (17.6 M methods) and C (6.3 M functions) corpora. Our results show that linear correlation between SLOC and CC is only moderate as a result of increasingly high variance. We further observe that aggregating CC and SLOC as well as performing a power transform improves the correlation. Our conclusion is that the observed linear correlation between CC and SLOC of Java methods or C functions is not strong enough to conclude that CC is redundant with SLOC. This conclusion contradicts earlier claims from literature but concurs with the widely accepted practice of measuring of CC next to SLOC. Copyright © 2015 John Wiley & Sons, Ltd.
Davy Landman, Alexander Serebrenik, Eric Bouwers, Jurgen J. Vinju
J. Softw. Evol. Process.4
2015 Optimizing hash-array mapped tries for fast and lean immutable JVM collections
abstract
The data structures under-pinning collection API (e.g. lists, sets, maps) in the standard libraries of programming languages are used intensively in many applications. The standard libraries of recent Java Virtual Machine languages, such as Clojure or Scala, contain scalable and well-performing immutable collection data structures that are implemented as Hash-Array Mapped Tries (HAMTs). HAMTs already feature efficient lookup, insert, and delete operations, however due to their tree-based nature their memory footprints and the runtime performance of iteration and equality checking lag behind array-based counterparts. This particularly prohibits their application in programs which process larger data sets. In this paper, we propose changes to the HAMT design that increase the overall performance of immutable sets and maps. The resulting general purpose design increases cache locality and features a canonical representation. It outperforms Scala’s and Clojure’s data structure implementations in terms of memory footprint and runtime efficiency of iteration (1.3–6.7x) and equality checking (3–25.4x).
Michael J. Steindorfer, Jurgen J. Vinju
OOPSLA2
2015 Reducing the Cost of Grammar-Based Testing Using Pattern Coverage
Cleverton Hentz, Jurgen J. Vinju, Anamaria Martins Moreira
ICTSS2
2015 OSSMETER: a software measurement platform for automatically analysing open source software projects
abstract
Deciding whether an open source software (OSS) project meets the required standards for adoption in terms of quality, maturity, activity of development and user support is not a straightforward process as it involves exploring various sources of information. Such sources include OSS source code repositories, communication channels such as newsgroups, forums, and mailing lists, as well as issue tracking systems. OSSMETER is an extensible and scalable platform that can monitor and incrementally analyse a large number of OSS projects. The results of this analysis can be used to assess various aspects of OSS projects, and to directly compare different OSS projects with each other.
Davide Di Ruscio, Dimitrios S. Kolovos, Ioannis Korkontzelos, Nicholas Drivalos Matragkas, Jurgen J. Vinju
ESEC/SIGSOFT FSE5
2015 Modular language implementation in Rascal - experience report
Bas Basten, Jeroen van den Bos, Mark Hills 0001, Paul Klint, Arnold Lankamp, Bert Lisser, Atze van der Ploeg, Tijs van der Storm, Jurgen J. Vinju
Sci. Comput. Program.9
2015 Towards multilingual programming environments
Tijs van der Storm, Jurgen J. Vinju
Sci. Comput. Program.2
2015 Preface
Jurgen J. Vinju
Sci. Comput. Program.1
2014 Code specialization for memory efficient hash tries (short paper)
abstract
The hash trie data structure is a common part in standard collection libraries of JVM programming languages such as Clojure and Scala. It enables fast immutable implementations of maps, sets, and vectors, but it requires considerably more memory than an equivalent array-based data structure. This hinders the scalability of functional programs and the further adoption of this otherwise attractive style of programming. In this paper we present a product family of hash tries. We generate Java source code to specialize them using knowledge of JVM object memory layout. The number of possible specializations is exponential. The optimization challenge is thus to find a minimal set of variants which lead to a maximal loss in memory footprint on any given data. Using a set of experiments we measured the distribution of internal tree node sizes in hash tries. We used the results as a guidance to decide which variants of the family to generate and which variants should be left to the generic implementation. A preliminary validating experiment on the implementation of sets and maps shows that this technique leads to a median decrease of 55% in memory footprint for maps (and 78% for sets), while still maintaining comparable performance. Our combination of data analysis and code specialization proved to be effective.
Michael J. Steindorfer, Jurgen J. Vinju
GPCE2
2014 Empirical Analysis of the Relationship between CC and SLOC in a Large Corpus of Java Methods
abstract
Measuring the internal quality of source code is one of the traditional goals of making software development into an engineering discipline. Cyclomatic Complexity (CC) is an often used source code quality metric, next to Source Lines of Code (SLOC). However, the use of the CC metric is challenged by the repeated claim that CC is redundant with respect to SLOC due to strong linear correlation. We test this claim by studying a corpus of 17.8M methods in 13K open-source Java projects. Our results show that direct linear correlation between SLOC and CC is only moderate, as caused by high variance. We observe that aggregating CC and SLOC over larger units of code improves the correlation, which explains reported results of strong linear correlation in literature. We suggest that the primary cause of correlation is the aggregation. Our conclusion is that there is no strong linear correlation between CC and SLOC of Java methods, so we do not conclude that CC is redundant with SLOC. This conclusion contradicts earlier claims from literature, but concurs with the widely accepted practice of measuring of CC next to SLOC.
Davy Landman, Alexander Serebrenik, Jurgen J. Vinju
ICSME3
2014 Static, lightweight includes resolution for PHP
abstract
Dynamic languages include a number of features that are challenging to model properly in static analysis tools. In PHP, one of these features is the include expression, where an arbitrary expression provides the path of the file to include at runtime. In this paper we present two complementary analyses for statically resolving PHP includes, one that works at the level of individual PHP files, and one targeting PHP programs possibly consisting of multiple scripts. To evaluate the effectiveness of these analyses we have applied the first to a corpus of 20 open-source systems, totaling more than 4.5 million lines of PHP, and the second to a number of programs from a subset of these systems. Our results show that, in many cases, includes can be resolved to a specific file or a small subset of possible files, enabling better IDE features and more advanced program analysis tools for PHP.
Mark Hills 0001, Paul Klint, Jurgen J. Vinju
ASE3
2013 Exploring the Limits of Domain Model Recovery
abstract
We are interested in re-engineering families of legacy applications towards using Domain-Specific Languages (DSLs). Is it worth to invest in harvesting domain knowledge from the source code of legacy applications? Reverse engineering domain knowledge from source code is sometimes considered very hard or even impossible. Is it also difficult for "modern legacy systems"? In this paper we select two open-source applications and answer the following research questions: which parts of the domain are implemented by the application, and how much can we manually recover from the source code? To explore these questions, we compare manually recovered domain models to a reference model extracted from domain literature, and measured precision and recall. The recovered models are accurate: they cover a significant part of the reference model and they do not contain much junk. We conclude that domain knowledge is recoverable from "modern legacy" code and therefore domain model recovery can be a valuable component of a domain re-engineering process.
Paul Klint, Davy Landman, Jurgen J. Vinju
ICSM3
2013 An empirical study of PHP feature usage: a static analysis perspective
abstract
PHP is one of the most popular languages for server-side application development. The language is highly dynamic, providing programmers with a large amount of flexibility. However, these dynamic features also have a cost, making it difficult to apply traditional static analysis techniques used in standard code analysis and transformation tools. As part of our work on creating analysis tools for PHP, we have conducted a study over a significant corpus of open-source PHP systems, looking at the sizes of actual PHP programs, which features of PHP are actually used, how often dynamic features appear, and how distributed these features are across the files that make up a PHP website. We have also looked at whether uses of these dynamic features are truly dynamic or are, in some cases, statically understandable, allowing us to identify specific patterns of use which can then be taken into account to build more precise tools. We believe this work will be of interest to creators of analysis tools for PHP, and that the methodology we present can be leveraged for other dynamic languages with similar features.
Mark Hills 0001, Paul Klint, Jurgen J. Vinju
ISSTA3
2013 Safe Specification of Operator Precedence Rules
Ali Afroozeh, Mark van den Brand, Adrian Johnstone, Elizabeth Scott, Jurgen J. Vinju
SLE5
2013 Preface to the special section on Language Descriptions Tools and Applications (LDTA'08 & '09)
Jurgen J. Vinju
Sci. Comput. Program.1
2012 What Does Control Flow Really Look Like? Eyeballing the Cyclomatic Complexity Metric
abstract
Assessing the understandability of source code remains an elusive yet highly desirable goal for software developers and their managers. While many metrics have been suggested and investigated empirically, the McCabe cyclomatic complexity metric (CC) - which is based on control flow complexity - seems to hold enduring fascination within both industry and the research community despite its known limitations. In this work, we introduce the ideas of Control Flow Patterns (CFPs) and Compressed Control Flow Patterns (CCFPs), which eliminate some repetitive structure from control flow graphs in order to emphasize high-entropy graphs. We examine eight well-known open source Java systems by grouping the CFPs of the methods into equivalence classes, and exploring the results. We observed several surprising outcomes: first, the number of unique CFPs is relatively low, second, CC often does not accurately reflect the intricacies of Java control flow, and third, methods with high CC often have very low entropy, suggesting that they may be relatively easy to understand. These findings challenge the widely-held belief that there is a clear-cut causal relationship between CC and understandability, and suggest that CC and similar measures need to be reconsidered as metrics for code understandability.
Jurgen J. Vinju, Michael W. Godfrey
SCAM1
2012 Meta-language Support for Type-Safe Access to External Resources
Mark Hills 0001, Paul Klint, Jurgen J. Vinju
SLE3
2011 Ambiguity Detection: Scaling to Scannerless
Hendrikus J. S. Basten, Paul Klint, Jurgen J. Vinju
SLE3
2011 Parse Forest Diagnostics with Dr. Ambiguity
Hendrikus J. S. Basten, Jurgen J. Vinju
SLE2
2011 RLSRunner: Linking Rascal with K for Program Analysis
Mark Hills 0001, Paul Klint, Jurgen J. Vinju
SLE3
2010 Prototyping a tool environment for run-time assertion checking in JML with communication histories
abstract
In this paper we present prototype tool-support for the runtime assertion checking of the Java Modeling Language (JML) extended with communication histories specified by attribute grammars. Our tool suite integrates Rascal, a meta programming language and ANTLR, a popular parser generator. Rascal instantiates a generic model of history updates for a given Java program annotated with history specifications. ANTLR is used for the actual evaluation of history assertions.
Frank S. de Boer, Stijn de Gouw, Jurgen J. Vinju
FTfJP@ECOOP3
2010 Mod4J: A Qualitative Case Study of Model-Driven Software Development
Vincent Lussenburg, Tijs van der Storm, Jurgen J. Vinju, Jos Warmer
MoDELS (2)3
2010 Automated generation of program translation and verification tools using annotated grammars
Diego Ordóñez Camacho, Kim Mens, Mark van den Brand, Jurgen J. Vinju
Sci. Comput. Program.4
2009 Faster Scannerless GLR Parsing
Giorgios Economopoulos, Paul Klint, Jurgen J. Vinju
CC3
2009 Accelerating the creation of customized, language-Specific IDEs in Eclipse
abstract
Full-featured integrated development environments have become critical to the adoption of new programming languages. Key to the success of these IDEs is the provision of services tailored to the languages. However, modern IDEs are large and complex, and the cost of constructing one from scratch can be prohibitive. Generators that work from language specifications reduce costs but produce environments that do not fully reflect distinctive language characteristics.
Philippe Charles, Robert M. Fuhrer, Stanley M. Sutton Jr., Evelyn Duesterwald, Jurgen J. Vinju
OOPSLA5
2009 RASCAL: A Domain Specific Language for Source Code Analysis and Manipulation
abstract
Many automated software engineering tools require tight integration of techniques for source code analysis and manipulation. State-of-the-art tools exist for both, but the domains have remained notoriously separate because different computational paradigms fit each domain best. This impedance mismatch hampers the development of new solutions because the desired functionality and scalability can only be achieved by repeated and ad hoc integration of different techniques. RASCAL is a domain-specific language that takes away most of this boilerplate by integrating source code analysis and manipulation at the conceptual, syntactic, semantic and technical level. We give an overview of the language and assess its merits by implementing a complex refactoring.
Paul Klint, Tijs van der Storm, Jurgen J. Vinju
SCAM3
2005 Generalized Type-Based Disambiguation of Meta Programs with Concrete Object Syntax
Martin Bravenboer, Rob Vermaas, Jurgen J. Vinju, Eelco Visser
GPCE3
2005 An Architecture for Context-Sensitive Formatting
abstract
We have taken a fixed set of formatting requirements for a Cobol system as spelled out in a standardization document, and applied generic formatting technology to implement them. It appeared that corporate conventions can dictate alignment that crosscuts the logical structure of a program, and can even dictate indentation that is dynamically computed from context information. We have developed and implemented a formatting architecture that allows arbitrary computational power for mapping language constructs to the Box language. The enabling feature is a hybrid format that merges Box expressions with parse trees. Much of the boilerplate part of formatting can still be automated by a default mapping to Box. Absolute tab stops, an important feature which is not found in many Box back-ends, is used extensively in our case study.
Mark van den Brand, A. Taeke Kooiker, Jurgen J. Vinju, Niels P. Veerman
ICSM3
2003 Environments for Term Rewriting Engines for Free!
Mark van den Brand, Pierre-Etienne Moreau, Jurgen J. Vinju
RTA3
2003 Term rewriting with traversal functions
abstract
Term rewriting is an appealing technique for performing program analysis and program transformation. Tree (term) traversal is frequently used but is not supported by standard term rewriting. We extend many-sorted, first-order term rewriting with traversal functions that automate tree traversal in a simple and type-safe way. Traversal functions can be bottom-up or top-down traversals and can either traverse all nodes in a tree or can stop the traversal at a certain depth as soon as a matching node is found. They can either define sort-preserving transformations or mappings to a fixed sort. We give small and somewhat larger examples of traversal functions and describe their operational semantics and implementation. An assessment of various applications and a discussion conclude the article.
Mark van den Brand, Paul Klint, Jurgen J. Vinju
ACM Trans. Softw. Eng. Methodol.3
2002 Disambiguation Filters for Scannerless Generalized LR Parsers
Mark van den Brand, Jeroen Scheerder, Jurgen J. Vinju, Eelco Visser
CC3
2001 The ASF+SDF Meta-environment: A Component-Based Language Development Environment
Mark van den Brand, Arie van Deursen, Jan Heering, Hayco de Jong, Merijn de Jonge, Tobias Kuipers, Paul Klint, Leon Moonen, Pieter A. Olivier, Jeroen Scheerder, Jurgen J. Vinju, Eelco Visser, Joost Visser 0001
CC11