VLDB 2026 Research / reviewers in the wild / expert
Vadim Zaytsev
dblp:91/1157
· DBLP profile ↗
42ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0001-7764-4224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 41 · 10 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scala Mixed-Paradigm Maintainability Metrics
Ivo Broekhof, Rinse van Hees, Nhat, Vadim Zaytsev |
SANER | 4 |
| 2026 | CPSLint: A Domain-Specific Language Providing Data Validation and Sanitisation for Industrial Cyber-Physical SystemsabstractIndustrial cyber-physical systems generate vast amounts of semi-structured time-series data that require careful preprocessing before they can be effectively used for machine learning applications such as fault detection and identification. Raw sensor datasets are often corrupted or incomplete, making it challenging to develop reliable solutions without proper data preparation and validation. In this paper, we introduce CPSLint, a domain-specific language for data validation and sanitisation. We present the design, implementation and evaluation of CPSLint, demonstrating its ability to automatically detect and correct common data corruption patterns while enabling non-programming domain experts to effectively prepare their data for analysis. We report evaluation results on a representative dataset, tracking memory consumption and CPU-time for sanitisation activities. Our approach offers several advantages over traditional methods, including reduced manual effort, guaranteed consistency and broader applicability across time-series datasets and projects. Uraz Odyurt, Ömer Sayilir, Mariëlle Stoelinga, Vadim Zaytsev |
SLE | 4 |
| 2026 | SLE as an Evolving Body of Knowledge: A 19-Year Comparative Analysis of Calls and ProceedingsabstractThe Software Language Engineering (SLE) community has accumulated nearly two decades of content. This content comprises not only papers as such, but also other tangible artefacts like conference descriptions, calls for papers, tools, repositories. Yet, the field's self-description (how it positions itself to potential contributors) and its accumulated output (what the community actually produces) are rarely studied together in a reproducible, fine-grained way. Vadim Zaytsev |
SLE | 1 |
| 2026 | Mining Frequent Structures in Conceptual ModelsabstractAbstract The challenge of using structured methods to represent knowledge is a well-documented issue in conceptual modeling and has been the focus of extensive research. It is widely recognized that adopting modeling patterns offers an effective structural approach for designing conceptual models. Patterns, in this context, refer to generalizable, recurring structures that provide solutions to common design problems. They significantly enhance both the understanding and improvement of the modeling process. Numerous experimental studies have demonstrated the undeniable value of using patterns in conceptual modeling. Despite this, the task of identifying patterns in conceptual models remains highly complex, and there is currently no systematic method for pattern discovery. To address this gap, this paper proposes a general approach for discovering frequent structures in conceptual modeling languages as a means to support pattern identification. Specifically, we focus on uncovering recurring structures that reflect the usage patterns of a given conceptual modeling language. As proof of concept, we implement our approach by focusing on two widely used conceptual modeling languages. This implementation includes an exploratory tool that integrates a frequent subgraph mining algorithm with graph manipulation techniques , such as graph visualization , graph clustering , and graph transformation . The tool processes multiple conceptual models and identifies recurrent structures based on various criteria. We validate the tool using two state-of-the-art curated datasets: one consisting of models encoded in OntoUML and the other in ArchiMate. The primary objective of our approach is to provide a support tool for language engineers. This tool can be used to identify both effective and ineffective modeling practices, enabling the refinement and evolution of conceptual modeling languages. Furthermore, it facilitates the reuse of accumulated expertise, ultimately supporting the creation of higher-quality models in a given language. Mattia Fumagalli, Tiago Prince Sales, Pedro Paulo F. Barcelos, Giovanni Micale, Philipp-Lorenz Glaser, Dominik Bork, Vadim Zaytsev, Diego Calvanese, Giancarlo Guizzardi |
Softw. Syst. Model. | 7 |
| 2025 | Refactoring Detection Across Languages: Leveraging Java-Trained Models for Detecting Class-Level Refactorings in Kotlin
Mohammad Mehdi Afkhami, Iman Hemati Moghadam, Vadim Zaytsev, MohammadHossein Ashoori, Hossein Bazmandegan |
SEAA | 3 |
| 2025 | Comparative Analysis of Pre-trained Code Language Models for Automated Program Repair via Code Infill GenerationabstractAutomated Program Repair (APR) has advanced significantly with the emergence of pre-trained Code Language Models (CLMs), enabling the generation of high-quality patches. However, selecting the most suitable CLM for APR remains challenging due to a range of factors, including accuracy, efficiency, and scalability, among others. These factors are interdependent and interact in complex ways, making the selection of a CLM for APR a multifaceted problem. Iman Hemati Moghadam, Oebele Lijzenga, Vadim Zaytsev |
GPCE | 3 |
| 2025 | CoCoCoLa: Code Completion Control LanguageabstractIn software development, the efficiency and accuracy of code completion systems are crucial for productivity and codebase discovery. From simple spell checkers to advanced AI-powered tools, there are more ways to complete your code than ever. This results in an explosion in the number of possible valid proposals, especially when working with today’s increasingly large codebases. Over the years, a lot of effort has been put into developing effective ranking systems to prioritise proposals with more potential. Yet developers still often struggle with an overwhelming number of suggestions, leading to reduced productivity and increased cognitive load. In this paper, instead of just performing completion by name, we propose CoCoCoLa — an alternative approach to give back the control over the presented proposals to the developer. By investigating the recorded code completion events, frequencies of desirable code elements’ properties were calculated to identify useful control factors. To avoid adding further complexity to the completion process, we propose a simple language, defined within the boundary of a valid identifier of the 50+ most popular software languages in 2024. This language allows developers to specify and filter for desired properties of the proposals. Nhat, Vadim Zaytsev |
GPCE | 2 |
| 2025 | Requirements for an Automated Assessment Tool for Learning Programming by DoingabstractAssessment of open-ended assignments such as programming projects is a complex and time-consuming task. When students learn to program, however, they benefit from receiving timely feedback, which requires an assessment of their current work. Our goal is to build a tool that assists in this process by partially automating the assessment of open-ended programming assignments. In this paper we discuss the requirements for this tool, based on interviews with teachers and other relevant stakeholders. Arthur Rump, Vadim Zaytsev, Angelika Mader |
ICST | 2 |
| 2025 | Extract, model, refine: improved modelling of program verification tools through data enrichmentabstractIn software engineering, models are used for many different things. In this paper, we focus on program verification, where we use models to reason about the correctness of systems. There are many different types of program verification techniques which provide different correctness guarantees. We investigate the domain of program verification tools and present a concise megamodel to distinguish these tools. We also present a data set of 400+ program verification tools. This data set includes the category of verification tool according to our megamodel, practical information such as input/output format, repository links and more. The practical information, such as last commit date, is kept up to date through the use of APIs. Moreover, part of the data extraction has been automated to make it easier to expand the data set. The categorisation enables software engineers to find suitable tools, investigate alternatives and compare tools. We also identify trends for each level in our megamodel. Our data set, publicly available at https://doi.org/10.4121/20347950, can be used by software engineers to enter the world of program verification and find a verification tool based on their requirements. This paper is an extended version of https://doi.org/10.1145/3550355.3552426. Sophie Lathouwers, Vadim Zaytsev |
Softw. Syst. Model. | 3 |
| 2024 | Visual Assurance in Refactoring Through Trace Equivalence of Control Flow GraphsabstractRefactoring large legacy codebases, even with industrial-strength tools, often leads to trust concerns with code owners, in particular when the codebase underwent significant changes. To provide more assurance to code owners, we integrate visual analytics into the refactoring process. This method involves transforming code into control flow graphs before and after refactoring, followed by trace equivalence analysis on these graphs. An innovative visualisation tool provides not only a comprehensive overview of the refactorings' impact across all files, but also offers detailed insights into the trace equivalence at individual file level. By presenting clear visual evidence of code equivalence before and after refactoring, our visualisation narrows the trust gap, offering refactoring experts and code owners a transparent and understandable view of the changes. We apply this visualisation on an industrial use case and discuss its effectiveness with refactoring experts. Céline Deknop, Johan Fabry, Kim Mens, Vadim Zaytsev |
SANER | 4 |
| 2024 | The Limits of the Identifiable: Challenges in Python Version Identification with Deep LearningabstractThe evolution of Python requires accurate version identification to facilitate compatibility and ongoing support. We extend previous work on deep learning models for Python version identification, where LSTM and CodeBERT achieved a 92% accuracy on short code snippets. We further expand these results to larger realistic files, utilising code segmentation techniques for varying input granularities. These techniques ranged from per-line analysis to larger code segments. Our findings show that while LSTM with CodeBERT embeddings maintained high accuracy on short snippets, performance significantly drops on longer segments, particularly in balancing information retention and misclassification risks. Notably, import-statement analysis, despite being the most intuitive indicator of version requirements, reached only a 30% accuracy. This exposes the limitations of our approach when encountering rare or user-defined modules. The findings expose the limitations of deep learning for language version identification, and suggest that alternative approaches may be necessary for high accuracy on larger datasets. Marcus Gerhold, Lola Solovyeva, Vadim Zaytsev |
SANER | 3 |
| 2024 | Extending Refactoring Detection to Kotlin: A Dataset and Comparative StudyabstractRefactoring, as one of the best practices in software development, has been also the centre of attention of much research. Particularly, a plethora of studies have been performed to understand the impact of refactorings on different dimensions of software development including software quality, program comprehension, fault-proneness, and non-functional requirements, among others. Among the employed approaches, analysing refactorings applied previously in real-world scenarios has been used by many researchers and proves to be a valuable way to delve deeper into the subject. The results of these research studies not only enhance our understanding of the advantages and potential drawbacks of refactorings but also guide us in developing more efficient automated refactoring tools based on how developers actually use refactorings in practice. However, the majority of studies in this regard have focused on refactorings applied in Java programs, and the other programming languages have received significantly less attention. In reality, the lack of comprehensive datasets of real-world applied refactorings makes it challenging for researchers to conduct comprehensive studies in programming languages other than Java. The primary obstacle can be the lack of automated tool support for identifying refactorings applied in programs implemented in other languages. To mitigate this limitation, we extended a previously available refactoring detection tool, Refdetect, to be able to identify refactorings applied in Kotlin programs. We conducted an experiment on 200 commits of 10 Kotlin repositories sourced on GitHub and compared the performance of our tool with an existing Kotlin refactoring detection tool called Ko T L I Nrmi Ne R. We found that our tool has a precision of 90% and a recall of 82 %, achieving an average F -score of 84 % which is 17 % better than the one achieved by KOTLINRMINER. We also provide the resulting dataset containing 2,043 true refactoring instances detected by at least one of Refdetect or Kotlinrminer and validated by one up to three refactoring experts. By releasing this initial dataset, we aim to address the existing gap in the availability of Kotlin refactoring datasets. Iman Hemati Moghadam, Mohammad Mehdi Afkhami, Parsa Kamalipour, Vadim Zaytsev |
SANER | 4 |
| 2024 | Deriving modernity signatures of codebases with static analysisabstractThis paper addresses the problem of determining the modernity of software systems by analysing the use of new language features and their adoption over time. We propose the concept of modernity signatures to estimate the age of a codebase, naturally adjusted for maintenance practices, such that the modernity of a regularly updated system would be above that of a more recently created one which neglects current features and best practices. This can provide insights into coding practices, codebase health and the evolution of software languages. We present case studies on PHP and Python code, demonstrating the effectiveness of modernity signatures in determining the age of a codebase without executing the code or performing extensive human inspection. The paper describes the technical implementation details of generating the modernity signature for both of these languages, including the use of existing tools like the PHP parser and Vermin. The findings suggest that modernity signatures can aid developers in many ways from choosing whether to use a system or how to approach its maintenance, to assessing usefulness of a language feature, thus providing a valuable tool for source code analysis and manipulation. Chris Admiraal, Wouter Van den Brink, Marcus Gerhold, Vadim Zaytsev, Cristian Zubcu |
J. Syst. Softw. | 4 |
| 2023 | Crossover: Towards Compiler-Enabled COBOL-C InteroperabilityabstractInteroperability across software languages is an important practical topic. In this paper, we take a deep dive into investigating and tackling the challenges involved with achieving interoperability between C and BabyCobol. The latter is a domain-specific language condensing challenges found in compiling legacy languages — borrowing directly from COBOL's data philosophy. Crossover, a compiler designed specifically to showcase the interoperability, exposes details of connecting a COBOL-like language with PICTURE clauses and re-entrant procedures, to C with primitive types and struct composites. Crossover features a C library for overcoming the differences between the data representations native to the respective languages. We illustrate the design process of Crossover and demonstrate its usage to provide a strategy to achieve interoperability between legacy and modern languages. The described process is aimed to be a blueprint for achievable interoperability between full-fledged COBOL and modern C-like programming languages. Mart van Assen, Manzi Aimé Ntagengerwa, Ömer Sayilir, Vadim Zaytsev |
GPCE | 4 |
| 2023 | Responsible Language Design
Vadim Zaytsev |
MODELSWARD | 1 |
| 2022 | Generating Customised Control Flow Graphs for Legacy Languages with Semi-ParsingabstractWe propose a tool and underlying technique that uses semi-parsing to extract control flow graphs from legacy source code (i.e., COBOL). Obtaining such control flow graphs is relevant in the industrial setting of legacy modernisation, to quickly demonstrate to code owners that modernisation engineers did not break their business logic. They need to be convinced that a migration did not affect the flow around critical parts of their code such as database accesses. Focusing on the control flow around embedded SQL queries and confirming that the code logic has been preserved improves customers' trust and satisfaction in the modernisation. Our proposed algorithm and approach uses fuzzy parsing as opposed to full parsing to parse mainly the control flow constructs, while delegating the full parsing of embedded languages like SQL to an external parser, and produces a control flow graph directly while skipping over most of the input in linear time. Such a fuzzy parser is easier to construct and adapt to particular languages and needs than a full parser with a visitor to elicit control flow. Comparisons are made of the fuzzy parser to an industrial-strength full parser. Céline Deknop, Johan Fabry, Kim Mens, Vadim Zaytsev |
ICSME | 4 |
| 2022 | Modelling program verification tools for software engineersabstractIn software engineering, models are used for many different things. In this paper, we focus on program verification, where we use models to reason about the correctness of systems. There are many different types of program verification techniques which provide different correctness guarantees. We investigate the domain of program verification tools, and present a concise megamodel to distinguish these tools. We also present a data set of almost 400 program verification tools. This data set includes the category of verification tool according to our megamodel, practical information such as input/output format, repository links, and more. The categorisation enables software engineers to find suitable tools, investigate similar alternatives and compare them. We also identify trends for each level in our megamodel based on the categorisation. Our data set, publicly available at https://doi.org/10.4121/20347950, can be used by software engineers to enter the world of program verification and find a verification tool based on their requirements. Sophie Lathouwers, Vadim Zaytsev |
MoDELS | 2 |
| 2022 | Deriving Modernity Signatures for PHP Systems with Static AnalysisabstractThe PHP language has undergone many changes in its syntax and grammar, with respect to both features the language has to offer as well as the distribution of language features used by programmers in their projects. We present a novel method of using grammar usage statistics to calculate a modernity signature for a PHP system, so that we can determine its age. The system will aid developers in choosing whether or not to execute or use a PHP system, without having to perform an extensive inspection. Wouter Van den Brink, Marcus Gerhold, Vadim Zaytsev |
SCAM | 3 |
| 2021 | There is more than one way to zen your PythonabstractThe popularity of Python can be at least partially attributed to the concept of pythonicity, loosely defined as a combination of good practices accepted within the community. Despite the popularity of both Python itself and the pythonicity of code written in it, this concept has not been studied that well, and the first attempts to define it formally are rather recent. In this paper, we take the next steps in exploring this topic by conducting an independent literature review in order to create a catalogue of pythonic idioms, reproduce the results of a recent paper on the usage of pythonic idioms, perform an external direct replication of it by reusing the same open source toolset and dataset, and extend the body of knowledge by also analysing how the use of pythonic idioms evolve over time in open source codebases. Aamir Farooq, Vadim Zaytsev |
SLE | 2 |
| 2021 | A Scalable Log Differencing Visualisation Applied to COBOL RefactoringabstractLarge code refactoring projects can consist of hundreds of refactoring rules that are applied iteratively to make code easier to maintain. Visualising the refactoring process can help engineers and stakeholders understand how chains of refactorings were applied and to gain more confidence in the produced result. An apparently suitable existing visualisation using log-based behavioural differencing suffers from scalability issues when applied to industrial-size cases. We propose an adapted visualisation tool that highlights those parts that really changed in-between iterations of a large refactoring process and collapses those parts that remain stable. We show that our alternative visualisation scales well on large logs of a process with many possible refactoring chains, of which significant parts are shared. Consequently, it allows engineers and stakeholders to quickly answer relevant questions about what happened during the refactoring process. Céline Deknop, Kim Mens, Alexandre Bergel, Johan Fabry, Vadim Zaytsev |
VISSOFT | 5 |
| 2020 | Improving a Software Modernisation Process by Differencing Migration Logs
Céline Deknop, Johan Fabry, Kim Mens, Vadim Zaytsev |
PROFES | 4 |
| 2020 | Software language engineers' worst nightmareabstractMany techniques in software language engineering get their first validation by being prototyped to work on one particular language such as Java, Scala, Scheme, or ML, or a subset of such a language. Claims of their generalisability, as well as discussion on potential threats to their external validity, are often based on authors' ad hoc understanding of the world outside their usual comfort zone. To facilitate and simplify such discussions by providing a solid measurable ground, we propose a language called BabyCobol, which was specifically designed to contain features that turn processing legacy programming languages such as COBOL, FORTRAN, PL/I, REXX, CLIST, and 4GLs (fourth generation languages), into such a challenge. The language is minimal by design so that it can help to quickly find weaknesses in frameworks making them inapplicable to dealing with legacy software. However, applying new techniques of software language engineering and reverse engineering to such a small language will not be too tedious and overwhelming. BabyCobol was designed in collaboration with industrial compiler developers by systematically traversing features of several second, third and fourth generation languages to identify the core culprits in making development of compiler for legacy languages difficult. Vadim Zaytsev |
SLE | 1 |
| 2019 | Mining Patterns in Source Code Using Tree Mining Algorithms
Hoang-Son Pham, Siegfried Nijssen, Kim Mens, Dario Di Nucci, Tim Molderez, Coen De Roover, Johan Fabry, Vadim Zaytsev |
DS | 8 |
| 2019 | Qualify First! A Large Scale Modernisation ReportabstractTypically in modernisation projects any concerns for code quality are silenced until the end of the migration, to simplify an already complex process. Yet, we claim from experience that prioritising quality above many other issues has many benefits. In this experience report, we discuss a modernisation project of mBank, a big Polish bank, where bad smell detection and elimination, automated testing and refactoring played a crucial rule, provided pay-offs early in the project, increased buy-in, and ensured maintainability of the end result. Leszek Wlodarski, Boris Pereira, Ivan Povazan, Johan Fabry, Vadim Zaytsev |
SANER | 5 |
| 2018 | An industrial case study in compiler testing (tool demo)abstractCompiler construction is one of the oldest areas of software engineering, yet despite its maturity it has underdeveloped sides such as compiler testing. There exist many disparate methods for testing parsers, optimisers and other components, but no unified methodology that consumable by practitioners from a book to be directly applied to fulfil their needs. Vadim Zaytsev |
SLE | 1 |
| 2017 | Parser generation by example for legacy pattern languagesabstractMost modern software languages enjoy relatively free and relaxed concrete syntax, with significant flexibility of formatting of the program/model/sheet text. Yet, in the dark legacy corners of software engineering there are still languages with a strict fixed column-based structure — the compromises of times long gone, attempting to combine some human readability with some ease of machine processing. In this paper, we consider an industrial case study for retirement of a legacy domain-specific language, completed under extreme circumstances: absolute lack of documentation, varying line structure, hierarchical blocks within one file, scalability demands for millions of lines of code, performance demands for manipulating tens of thousands multi-megabyte files, etc. However, the regularity of the language allowed to infer its structure from the available examples, automatically, and produce highly efficient parsers for it. Vadim Zaytsev |
GPCE | 1 |
| 2017 | Language Design with IntentabstractSoftware languages have always been an essential component of model-driven engineering. Their importance and popularity has been on the rise thanks to language workbenches, language-oriented development and other methodologies that enable us to quickly and easily create new languages specific for each domain. Unfortunately, language design is largely a form of art and has resisted most attempts to turn it into a form of science or engineering. In this paper we borrow concepts, techniques and principles from the domain of persuasive technology, or wider yet, design with intent --- which was developed as a way to influence users behaviour for social and environmental benefit. Similarly, we claim, software language designers can make conscious choices in order to influence the behaviour of language users. The paper describes a process of extracting design components from 24 books of eight categories (dragon books, parsing techniques, compiler construction, compiler design, language implementation, language documentation, programming languages, software languages), as well as from the original set of Design with Intent cards and papers on DSL design. The resulting language design card toolkit can be used by DSL designers to cover important design decisions and make them with more confidence. Vadim Zaytsev |
MoDELS | 1 |
| 2017 | Towards a taxonomy of grammar smellsabstractAny grammar engineer can tell a good grammar from a bad one, but there is no commonly accepted taxonomy of indicators of required grammar refactorings. One of the consequences of this lack of general smell taxonomy is the scarcity of tools to assess and improve the quality of grammars. By combining two lines of research — on smell detection and on grammar transformation — we have assembled a taxonomy of smells in grammars. As a pilot case, the detectors for identified smells were implemented for grammars in a broad sense and applied to the 641 grammars of the Grammar Zoo. Mats Stijlaart, Vadim Zaytsev |
SLE | 2 |
| 2016 | The A?B*A Pattern: Undoing Style in CSS and Refactoring Opportunities It PresentsabstractCascading Style Sheets (CSS) is a language widely used in contemporary web applications for defining the presentation semantics of web documents. Despite its relatively simple syntax, the language has a number of complex features like inheritance, cascading and specificity, which make CSS code challenging to understand and maintain. It has been noted in prior research that CSS code is prone to contain code smells which indicate design weaknesses and maintainability issues. In this paper we focus on one of those code smells called undoing style. It happens when a property is set to a value A, then overridden to another value B, possibly multiple times, and then set back to the original value of A. We refer to this pattern as the A?B*A pattern. We propose a technique that detects undoing style in CSS code and recommends refactoring opportunities to eliminate instances of undoing style while preserving the semantics of the web application. We evaluate our technique on 41 real-world web applications, and outline a proof of correctness for our refactoring. Our findings show that undoing style is quite prominent in CSS code. Additionally, there are many refactorings that can be applied while hardly introducing any errors. Leonard Punt, Sjoerd Visscher, Vadim Zaytsev |
ICSME | 3 |
| 2016 | A Tool for Detecting and Refactoring the A?B*A Pattern in CSSabstractThis tool detects undoing style in CSS code and is able to apply refactoring opportunities to eliminate a subset of these instances of undoing style, while preserving the semantics of the web application. Leonard Punt, Sjoerd Visscher, Vadim Zaytsev |
ICSME | 3 |
| 2016 | Experimental Data for the A?B*A Pattern in CSS: Inputs and OutputsabstractThis dataset is used to detect undoing style in CSS code. In total, this dataset contains 41 subjects. Each subject has its own folder, which contains the captured states, a states.html file, is used to load all captured states in one document, and a folder called results, which contains the detected undoing styles, the refactored style sheets and the detected semantic changes. The file states.html was used as an input for our detection tool. Leonard Punt, Sjoerd Visscher, Vadim Zaytsev |
ICSME | 3 |
| 2016 | Raincode assembler compiler (tool demo)
Volodymyr Blagodarov, Yves Jaradin, Vadim Zaytsev |
SLE | 3 |
| 2016 | Language design and implementation for the domain of coding conventions
Boryana Goncharenko, Vadim Zaytsev |
SLE | 2 |
| 2016 | Software Language Identification with Natural Language ClassifiersabstractSoftware language identification techniques are applicable to many situations from universal IDE support to legacy code analysis. Most widely used heuristics are based on software artefact metadata such as file extensions or on grammar-based text analysis such as keyword search. In this paper we propose to use statistical language models from the natural language processing field such as n-grams, skip-grams, multinominal naïve Bayes and normalised compression distance. Our preliminary experiments show that some of these models used as classifiers can achieve high precision and recall and can be used to properly identify language families, languages and even deal with embedded code fragments. Juriaan Kennedy van Dam, Vadim Zaytsev |
SANER | 2 |
| 2015 | Grammar Zoo: A corpus of experimental grammarwareabstractIn this paper we describe composition of a corpus of grammars in a broad sense in order to enable reuse of knowledge accumulated in the field of grammarware engineering. The Grammar Zoo displays the results of grammar hunting for big grammars of mainstream languages, as well as collecting grammars of smaller DSLs and extracting grammatical knowledge from other places. It is already operational and publicly supplies its users with grammars that have been recovered from different sources of grammar knowledge, varying from official language standards to community-created wiki pages. We summarise recent achievements in the discipline of grammarware engineering, that made the creation of such a corpus possible. We also describe in detail the technology that is used to build and extend such a corpus. The current contents of the Grammar Zoo are listed, as well as some possible future uses for them. Vadim Zaytsev |
Sci. Comput. Program. | 1 |
| 2014 | Parsing in a Broad Sense
Vadim Zaytsev, Anya Helene Bagge |
MoDELS | 1 |
| 2013 | Micropatterns in Grammars
Vadim Zaytsev |
SLE | 1 |
| 2011 | Comparison of Context-Free Grammars Based on Parsing Generated Test Data
Bernd Fischer 0002, Ralf Lämmel, Vadim Zaytsev |
SLE | 3 |
| 2011 | Recovering grammar relationships for the Java Language Specification
Ralf Lämmel, Vadim Zaytsev |
Softw. Qual. J. | 2 |
| 2010 | A Unified Format for Language Documents
Vadim Zaytsev, Ralf Lämmel |
SLE | 1 |
| 2009 | An Introduction to Grammar Convergence
Ralf Lämmel, Vadim Zaytsev |
IFM | 2 |
| 2009 | Recovering Grammar Relationships for the Java Language SpecificationabstractWe describe a completed effort to recover the relationships between all the grammars that occur in the different versions of the Java Language Specification (JLS). The relationships are represented as grammar transformations that capture all accidental or intended differences between the JLS grammars. This process is mechanized and it is driven by simple measures of nominal or structural differences between any pair of grammars involved. Our work suggests a form of consistency management for the JLS in particular, and language specifications in general. Ralf Lämmel, Vadim Zaytsev |
SCAM | 2 |