VLDB 2026 Research / reviewers in the wild / expert
Zoltán Porkoláb
dblp:65/2241
· DBLP profile ↗
21ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-6819-0224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Overlord: A C++ Overloading InspectorabstractFunction overloading is a well-known technique available in most major programming languages, including$\mathrm{C}^{++}$, to facilitate compile-time polymorphism. Due to the complex requirements imposed by the$\mathbf{C}^{++}$Language Standard, the behaviour of overload resolution can be surprising even for experienced$\mathbf{C}^{++}$developers. This is exacerbated by the fact that if an unexpected function is selected in an overloaded situation, most compilers do not explain why. In addition, an overabundance of candidates can be a compilation time bottleneck. Despite its widespread usage, there are not many tools to help programmers analyse overload usage in a software project. We developed an overloading inspector tool, Overlord, based on the open-source LLVM/Clang Compiler Infrastructure. With Overlord developers can list the possible candidate functions for a call site, and obtain step-by-step reasoning about the candidate selection process. Additionally to its comprehension functionality the tool also provides profiling data on overload resolution times to help library authors streamline the set of available overloads for improving the compilation performance. Botond István Horváth, Richárd Szalay, Zoltán Porkoláb |
ICPC | 3 |
| 2025 | Refactoring to Standard C++20 ModulesabstractABSTRACT Good component‐based design for software projects is a desired property both for development and maintenance. The C++ programming language inherited the “translation unit” model from C, where every source file is individually compiled with no knowledge about other parts of the project. This model has several drawbacks, and C++20 Modules is the Standard's answer for them. Moreover, Modules allows a cleaner encapsulation of concern. This paper investigates a semi‐automatic modularization method to refactor existing C++ projects. Our approach uses dependency analysis and clustering to organize elements of an existing project into modules, without domain‐specific information. Based on our study of two medium‐size open‐source projects from disjoint domains and vastly distinct architecture, upgrading existing software systems to the new Modules feature is limited by the existing design of the project's architecture. To fully facilitate the use of Modules in a project, it is likely that both project‐internal and user‐facing interfaces must be changed. Richárd Szalay, Zoltán Porkoláb |
J. Softw. Evol. Process. | 2 |
| 2022 | An Evaluation of General-Purpose Static Analysis Tools on C/C++ Test CodeabstractIn recent years, maintaining test code quality has gained more attention due to increased automation and the growing focus on issues caused during this process.Test code may become long and complex, but maintaining its quality is mostly a manual process, that may not scale in big software projects. Moreover, bugs in test code may give a false impression about the correctness or performance of the production code. Static program analysis (SPA) tools are being used to maintain the quality of software projects nowadays. However, these tools are either not used to analyse test code, or any analysis results on the test code are suppressed.This is especially true since SPA tools are not tailored to generate precise warnings on test code. This paper investigates the use of SPA on test code by employing three state-of-the-art general-purpose static analysers on a curated set of projects used in the industry and a random sample of relatively popular and large open-source C/C++ projects. We have found a number of built-in code checking modules that can detect quality issues in the test code. However, these checkers need some tailoring to obtain relevant results. We observed design choices in test frameworks that raise noisy warnings in analysers and propose a set of augmentations to the checkers or the analysis framework to obtain precise warnings from static analysers. Jean Malm, Eduard Paul Enoiu, Abu Naser Masud, Björn Lisper, Zoltán Porkoláb, Sigrid Eldh |
SEAA | 5 |
| 2022 | Semi-Automatic Refactoring to C++20 Modules: A Semi-Success StoryabstractThe component-based design of software projects is a desired property both for development and ease of code comprehension. Programming languages have long allowed component-based development (e.g., Java packages, Python modules); however, other languages, especially C and C++, had stuck to the “translation unit” model where every source file is individually compiled. The Modules system of C++20 was expected to allow cleaner encapsulation of concern. In this paper, we investigate the effort of a (semi-)automatic modularisation of existing C++ projects. Based on our investigation, upgrading existing software systems to the new Modules feature is extremely hard due to coupling issues arising from necessarily legacy design. Implementing real transition requires a significant redesign of both project-internal and user-facing programming interfaces. Richárd Szalay, Zoltán Porkoláb |
SCAM | 2 |
| 2022 | Flexible semi-automatic support for type migration of primitives for C/C++ programsabstractType systems are crucial tools in the hands of developers to ensure an increased level of soundness to their programs, make them safer, and guard against bugs. However, in practice, the type system is not always used to its full capability, and trade-offs are made. The effects range from hindered code comprehension and wasted development effort to financial damages and even the risk of loss of life. Various widely used programming languages, such as C++, Java, Python, and Rust, are yet to implement constrained types as a language feature. While the usage of user-defined types is common in modern languages that support such elements, developers often resort to having their variables use the most common, fundamental, built-in or library types, such as int or string, and not encode invariants into the type system. In this paper, we describe a flexible, incremental, semi-automated type migration approach that allows transitioning from the use of fundamental coarse types to strong types. Previously, well-scaling type migration methods were restricted and required the existence of an already well-defined destination type. In our case, the new strong type to be created is not yet defined, but the required interface is discovered via static analysis. As refactoring tools would cease to function if the code is changed textually to a version that refers undefined symbols, a more granular approach was needed. In addition, our proposed method allows us to discover the mixing of conceptually distinct types during development while the original program continues to function. Richárd Szalay, Zoltán Porkoláb |
SANER | 2 |
| 2021 | Unambiguity of Python Language Elements for Static AnalysisabstractStatic analysis is a technique for gathering some meaningful information during compilation-time and using it for further processing. This technique is applied in bug finding, code comprehension and in many other areas. However, static analysis gives a static view of the source and provides limited information about the dynamic behavior of a program. This limitation is generally considered to be a major barrier for dynamically typed languages, like Python, since in many situations it is hard to derive the most basic properties of variables or functions, sometimes we can’t even find the location of their definition. In this paper we present our experimental results that show this is not necessary a hard barrier in practical industrial projects. We analyzed open-source Python libraries, including the Python Standard Library itself as part of the CodeCompass code comprehension framework, and found that in a high number of the cases it was possible to decide the mentioned important features in unambiguous way. The consistent coding conventions make it possible to reduce unambiguity in static analysis of dynamic languages. Bence Nagy, Tibor Brunner, Zoltán Porkoláb |
SCAM | 3 |
| 2021 | Practical heuristics to improve precision for erroneous function argument swapping detection in C and C++abstractArgument selection defects, in which the programmer chooses the wrong argument to pass to a parameter from a potential set of arguments in a function call, is a widely investigated problem. The compiler can detect such misuse of arguments only through the argument and parameter type for statically typed programming languages. When adjacent parameters have the same type or can be converted between one another, a swapped or out of order call will not be diagnosed by compilers. Related research is usually confined to exact type equivalence, often ignoring potential implicit or explicit conversions. However, in current mainstream languages, like C++, built-in conversions between numerics and user-defined conversions may significantly increase the number of mistakes to go unnoticed. We investigated the situation for C and C++ languages where developers can define functions with multiple adjacent parameters that allow arguments to pass in the wrong order. When implicit conversions – such as parameter pairs of types – are taken into account, the number of mistake-prone functions markedly increases compared to only strict type equivalence. We analysed a sample of projects and categorised the offending parameter types. The empirical results should further encourage the language and library development community to emphasise the importance of strong typing and to restrict the proliferation of implicit conversions. However, the analysis produces a hard to consume amount of diagnostics for existing projects, and there are always cases that match the analysis rule but cannot be “fixed”. As such, further heuristics are needed to allow developers to refactor effectively based on the analysis results. We devised such heuristics, measured their expressive power, and found that several simple heuristics greatly help highlight the more problematic cases. Richárd Szalay, Ábel Sinkovics, Zoltán Porkoláb |
J. Syst. Softw. | 3 |
| 2020 | The Role of Implicit Conversions in Erroneous Function Argument Swapping in C++abstractArgument selection defects, in which the programmer has chosen the wrong argument to a function call is a widely investigated problem. The compiler can detect such misuse of arguments based on the argument and parameter type in case of statically typed programming languages. When adjacent parameters have the same type, or they can be converted between one another, the potential error will not be diagnosed. Related research is usually confined to exact type equivalence, often ignoring potential implicit or explicit conversions. However, in current mainstream languages, like C++, built-in conversions between numerics and user-defined conversions may significantly increase the number of mistakes to go unnoticed. We investigated the situation for C and C++ languages where functions are defined with multiple adjacent parameters that allow arguments to pass in the wrong order. When implicit conversions are taken into account, the number of mistake-prone function declarations significantly increases compared to strict type equivalence. We analysed the outcome and categorised the offending parameter types. The empirical results should further encourage the language and library development community to emphasise the importance of strong typing and the restriction of implicit conversion. Richárd Szalay, Ábel Sinkovics, Zoltán Porkoláb |
SCAM | 3 |
| 2018 | The codecompass comprehension frameworkabstractCodeCompass is an open source LLVM/Clang based tool developed by Ericsson Ltd. and the Eötvös Loránd University, Budapest to help understanding large legacy software systems. Based on the LLVM/Clang compiler infrastructure, CodeCompass gives exact information on complex C/C++ language elements like overloading, inheritance, the usage of variables and types, possible uses of function pointers and the virtual functions - features that various existing tools support only partially. Steensgaard's and Andersen's pointer analysis algorithm are used to compute and visualize the use of pointers/references. The wide range of interactive visualizations extends further than the usual class and function call diagrams; architectural, component and interface diagrams are a few of the implemented graphs. To make comprehension more extensive, CodeCompass is not restricted to the source code. It also utilizes build information to explore the system architecture as well as version control information e.g. git commit history and blame view. Clang based static analysis results are also integrated to CodeCompass. Although the tool focuses mainly on C and C++, it also supports Java and Python languages. Zoltán Porkoláb, Tibor Brunner |
ICPC | 1 |
| 2018 | Codecompass: an open software comprehension framework for industrial usageabstractCodeCompass is an open source LLVM/Clang-based tool developed by Ericsson Ltd. and Eötvös Loránd University, Budapest to help the understanding of large legacy software systems. Based on the LLVM/Clang compiler infrastructure, CodeCompass gives exact information on complex C/C++ language elements like overloading, inheritance, the usage of variables and types, possible uses of function pointers and virtual functions - features that various existing tools support only partially. Steensgaard's and Andersen's pointer analysis algorithms are used to compute and visualize the use of pointers/references. The wide range of interactive visualizations extends further than the usual class and function call diagrams; architectural, component and interface diagrams are a few of the implemented graphs. To make comprehension more extensive, CodeCompass also utilizes build information to explore the system architecture as well as version control information. Zoltán Porkoláb, Tibor Brunner, Dániel Krupp, Márton Csordás |
ICPC | 1 |
| 2018 | Selective friends in C++abstractSummary There is a strong prejudice against the friendship access control mechanism in C++. People claim that friendship breaks the encapsulation, reflects bad design, and creates too strong coupling. However, friends appear even in the most carefully designed systems, and if it is used judiciously (like using the attorney‐client idiom), they may be better choice than widening the public interface of the class. In this paper, we investigate how the friendship mechanism is used in C++ programs. We have made measurements on several open source projects to understand the current use of friends. Our results show various holes and errors in friend usage, like friend functions accessing only public members or not accessing members at all or the class, which declare friends has no private members at all. The results also show that friend functions actually use only a low percentage of the private members they were granted to access, which is a source of errors. These results have motivated us to propose a selective friend language construct for C++, which can restrict friendship only to well‐defined members. Such a new language element may decrease the degradation of encapsulation and significantly increase the diagnostic capacity of the compiler. We have created a proof‐of‐concept implementation based on the LLVM/Clang compiler infrastructure to show that such constructs can be established with a minimal syntactical and compilation overhead. Gábor Márton, Zoltán Porkoláb |
Softw. Pract. Exp. | 2 |
| 2017 | Towards Better Symbol Resolution for C/C++ Programs: A Cluster-Based SolutionabstractResolving symbol references is an important part of many application areas from development environments to various static analyser tools, especially when it is used for code comprehension purposes. Different occurrences of the same program elements, like function definitions and their call sites, variable declarations and their usage, or type definitions and their applications should be connected. In case of the C++ programming language, the most current tools use mangled names to correlate symbols, e.g. when implementing actions like "go to definition" or "list all references". However, for large projects, where multiple binaries are created, symbol resolution based on mangled names can be, and usually is, ambiguous. This leads to inaccurate behaviour even in major development tools. In this paper we explore the reason of this ambiguity, and propose our clustering algorithm based on essential build information to improve the accuracy of symbol resolution. We implemented our method as part of the CodeCompass open source code comprehension tool and measured its efficiency. Richárd Szalay, Zoltán Porkoláb, Dániel Krupp |
SCAM | 2 |
| 2013 | Implementing monads for C++ template metaprograms
Ábel Sinkovics, Zoltán Porkoláb |
Sci. Comput. Program. | 2 |
| 2011 | Extension of Iterator Traits in the C++ Standard Template Library
Norbert Pataki, Zoltán Porkoláb |
FedCSIS | 2 |
| 2011 | Type-preserving heap profiler for C++abstractMemory profilers are essential tools to understand the dynamic behaviour of complex modern programs. They help to reveal memory handling details: the wheres, the whens and the whats of memory allocations. Most heap profilers provide sufficient information about which part of the source code is responsible for the memory allocations by showing us the relevant call stacks. The sequence of allocations inform us about their order. However, in case of some strongly typed programming languages, like C++, the question what has been allocated is not trivial. Reporting the actual allocation size gives minimal or no information about the structure or type of the allocated objects. Though this information can be retrieved from the location and time of allocation, it cannot be easily automated, if at all. Therefore in large software systems programmers do not have an overall picture of which data structures are responsible for bottlenecks and have too few clues for pinpointing enhancement possibilities. In this paper we present a type-preserving heap profiler for C++. On top of the usual heap profiler features our allocation entries, including those of template constructs, contain exact type information about the allocated objects. We can extract information on individual memory operations as well as supply aggregated overview. Having such a type information in hand programmers can identify critical classes more easily and can perform optimizations based on evidence rather than speculations. József Mihalicza, Zoltán Porkoláb, Abel Gabor |
ICSM | 2 |
| 2010 | Domain-specific language integration with compile-time parser generator libraryabstractSmooth integration of domain-specific languages into a general purpose host language requires absorbing of domain code written in arbitrary syntax. The integration should cause minimal syntactical and semantic overhead and introduce minimal dependency on external tools. In this paper we discuss a DSL integration technique for the C++ programming language. The solution is based on compile-time parsing of the DSL code. The parser generator is a C++ template metaprogram reimplementation of a runtime Haskell parser generator library. The full parsing phase is executed when the host program is compiled. The library uses only standard C++ language features, thus our solution is highly portable. As a demonstration of the power of this approach, we present a highly efficient and type-safe version of printf and the way it can be constructed using our library. Despite the well known syntactical difficulties of C++ template metaprograms, building embedded languages using our library leads to self-documenting C++ source code. Zoltán Porkoláb, Ábel Sinkovics |
GPCE | 1 |
| 2010 | Visualization of C++ Template MetaprogramsabstractTemplate metaprograms have become an essential part of today's C++ programs: with proper template definitions we can force the C++ compiler to execute algorithms at compilation time. Among the application areas of template metaprograms are the expression templates, static interface checking, code optimization with adaptation, language embedding and active libraries. Despite all of its already proven benefits and numerous successful applications there are surprisingly few tools for creating, supporting, and analyzing C++ template metaprograms. As metaprograms are executed at compilation time they are even harder to understand. In this paper we present a code visualization tool, which is utilizing Tem plight, our previously developed C++ template metaprogram debugger. Using the tool it is possible to visualize the instantiation chain of C++ templates and follow the execution of metaprograms. Various presentation layers, filtering of template instances and step-by-step replay of the instantiations are supported. Our tool can help to test, optimize, maintain C++ template metaprograms, and can enhance their acceptance in the software industry. Zoltan Borok-Nagy, Viktor Majer, József Mihalicza, Norbert Pataki, Zoltán Porkoláb |
SCAM | 5 |
| 2008 | HypereiDoc - An XML Based Framework Supporting Cooperative Text Editions
Péter Bauer, Zsolt Hernáth, Zoltán Horváth, Gyula Mayer, Zsolt Parragi, Zoltán Porkoláb, Zsolt Sztupák |
ADBIS | 6 |
| 2006 | Debugging C++ template metaprogramsabstractTemplate metaprogramming is an emerging new direction in C++ programming for executing algorithms in compilation time. Despite all of its already proven benefits and numerous successful applications, it is yet to be accepted in industrial projects. One reason is the lack of professional software tools supporting the development of template metaprograms. A strong analogue exists between traditional runtime programs and compile-time metaprograms. This connection presents the possibility for creating development tools similar to those already used when writing runtime programs. This paper introduces Templight, a debugging framework that reveals the steps executed by the compiler during the compilation of C++ programs with templates. Templight's features include following the instantiation chain, setting breakpoints, and inspecting metaprogram information. This framework aims to take a step forward to help template metaprogramming become more accepted in the software industry. Zoltán Porkoláb, József Mihalicza, Ádám Sipos |
GPCE | 1 |
| 2004 | Towards a General Template Introspection Library
István Zólyomi, Zoltán Porkoláb |
GPCE | 2 |
| 2003 | An Extension to the Subtype Relationship in C++ Implemented with Template Metaprogramming
István Zólyomi, Zoltán Porkoláb, Tamás Kozsik |
GPCE | 2 |