Richárd Szalay

dblp:206/3339 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-5684-5158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 6 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Overlord: A C++ Overloading Inspector
abstract
Function overloading is a well-known technique available in most major programming languages, including$\mathrm{C}^{++}$, to facilitate compile-time polymorphism. Due to the complex requirements imposed by the$\mathbf{C}^{++}$Language Standard, the behaviour of overload resolution can be surprising even for experienced$\mathbf{C}^{++}$developers. This is exacerbated by the fact that if an unexpected function is selected in an overloaded situation, most compilers do not explain why. In addition, an overabundance of candidates can be a compilation time bottleneck. Despite its widespread usage, there are not many tools to help programmers analyse overload usage in a software project. We developed an overloading inspector tool, Overlord, based on the open-source LLVM/Clang Compiler Infrastructure. With Overlord developers can list the possible candidate functions for a call site, and obtain step-by-step reasoning about the candidate selection process. Additionally to its comprehension functionality the tool also provides profiling data on overload resolution times to help library authors streamline the set of available overloads for improving the compilation performance.
Botond István Horváth, Richárd Szalay, Zoltán Porkoláb
ICPC2
2025 Refactoring to Standard C++20 Modules
abstract
ABSTRACT Good component‐based design for software projects is a desired property both for development and maintenance. The C++ programming language inherited the “translation unit” model from C, where every source file is individually compiled with no knowledge about other parts of the project. This model has several drawbacks, and C++20 Modules is the Standard's answer for them. Moreover, Modules allows a cleaner encapsulation of concern. This paper investigates a semi‐automatic modularization method to refactor existing C++ projects. Our approach uses dependency analysis and clustering to organize elements of an existing project into modules, without domain‐specific information. Based on our study of two medium‐size open‐source projects from disjoint domains and vastly distinct architecture, upgrading existing software systems to the new Modules feature is limited by the existing design of the project's architecture. To fully facilitate the use of Modules in a project, it is likely that both project‐internal and user‐facing interfaces must be changed.
Richárd Szalay, Zoltán Porkoláb
J. Softw. Evol. Process.1
2022 Semi-Automatic Refactoring to C++20 Modules: A Semi-Success Story
abstract
The component-based design of software projects is a desired property both for development and ease of code comprehension. Programming languages have long allowed component-based development (e.g., Java packages, Python modules); however, other languages, especially C and C++, had stuck to the “translation unit” model where every source file is individually compiled. The Modules system of C++20 was expected to allow cleaner encapsulation of concern. In this paper, we investigate the effort of a (semi-)automatic modularisation of existing C++ projects. Based on our investigation, upgrading existing software systems to the new Modules feature is extremely hard due to coupling issues arising from necessarily legacy design. Implementing real transition requires a significant redesign of both project-internal and user-facing programming interfaces.
Richárd Szalay, Zoltán Porkoláb
SCAM1
2022 Flexible semi-automatic support for type migration of primitives for C/C++ programs
abstract
Type systems are crucial tools in the hands of developers to ensure an increased level of soundness to their programs, make them safer, and guard against bugs. However, in practice, the type system is not always used to its full capability, and trade-offs are made. The effects range from hindered code comprehension and wasted development effort to financial damages and even the risk of loss of life. Various widely used programming languages, such as C++, Java, Python, and Rust, are yet to implement constrained types as a language feature. While the usage of user-defined types is common in modern languages that support such elements, developers often resort to having their variables use the most common, fundamental, built-in or library types, such as int or string, and not encode invariants into the type system. In this paper, we describe a flexible, incremental, semi-automated type migration approach that allows transitioning from the use of fundamental coarse types to strong types. Previously, well-scaling type migration methods were restricted and required the existence of an already well-defined destination type. In our case, the new strong type to be created is not yet defined, but the required interface is discovered via static analysis. As refactoring tools would cease to function if the code is changed textually to a version that refers undefined symbols, a more granular approach was needed. In addition, our proposed method allows us to discover the mixing of conceptually distinct types during development while the original program continues to function.
Richárd Szalay, Zoltán Porkoláb
SANER1
2021 Practical heuristics to improve precision for erroneous function argument swapping detection in C and C++
abstract
Argument selection defects, in which the programmer chooses the wrong argument to pass to a parameter from a potential set of arguments in a function call, is a widely investigated problem. The compiler can detect such misuse of arguments only through the argument and parameter type for statically typed programming languages. When adjacent parameters have the same type or can be converted between one another, a swapped or out of order call will not be diagnosed by compilers. Related research is usually confined to exact type equivalence, often ignoring potential implicit or explicit conversions. However, in current mainstream languages, like C++, built-in conversions between numerics and user-defined conversions may significantly increase the number of mistakes to go unnoticed. We investigated the situation for C and C++ languages where developers can define functions with multiple adjacent parameters that allow arguments to pass in the wrong order. When implicit conversions – such as parameter pairs of types – are taken into account, the number of mistake-prone functions markedly increases compared to only strict type equivalence. We analysed a sample of projects and categorised the offending parameter types. The empirical results should further encourage the language and library development community to emphasise the importance of strong typing and to restrict the proliferation of implicit conversions. However, the analysis produces a hard to consume amount of diagnostics for existing projects, and there are always cases that match the analysis rule but cannot be “fixed”. As such, further heuristics are needed to allow developers to refactor effectively based on the analysis results. We devised such heuristics, measured their expressive power, and found that several simple heuristics greatly help highlight the more problematic cases.
Richárd Szalay, Ábel Sinkovics, Zoltán Porkoláb
J. Syst. Softw.1
2020 The Role of Implicit Conversions in Erroneous Function Argument Swapping in C++
abstract
Argument selection defects, in which the programmer has chosen the wrong argument to a function call is a widely investigated problem. The compiler can detect such misuse of arguments based on the argument and parameter type in case of statically typed programming languages. When adjacent parameters have the same type, or they can be converted between one another, the potential error will not be diagnosed. Related research is usually confined to exact type equivalence, often ignoring potential implicit or explicit conversions. However, in current mainstream languages, like C++, built-in conversions between numerics and user-defined conversions may significantly increase the number of mistakes to go unnoticed. We investigated the situation for C and C++ languages where functions are defined with multiple adjacent parameters that allow arguments to pass in the wrong order. When implicit conversions are taken into account, the number of mistake-prone function declarations significantly increases compared to strict type equivalence. We analysed the outcome and categorised the offending parameter types. The empirical results should further encourage the language and library development community to emphasise the importance of strong typing and the restriction of implicit conversion.
Richárd Szalay, Ábel Sinkovics, Zoltán Porkoláb
SCAM1
2017 Towards Better Symbol Resolution for C/C++ Programs: A Cluster-Based Solution
abstract
Resolving symbol references is an important part of many application areas from development environments to various static analyser tools, especially when it is used for code comprehension purposes. Different occurrences of the same program elements, like function definitions and their call sites, variable declarations and their usage, or type definitions and their applications should be connected. In case of the C++ programming language, the most current tools use mangled names to correlate symbols, e.g. when implementing actions like "go to definition" or "list all references". However, for large projects, where multiple binaries are created, symbol resolution based on mangled names can be, and usually is, ambiguous. This leads to inaccurate behaviour even in major development tools. In this paper we explore the reason of this ambiguity, and propose our clustering algorithm based on essential build information to improve the accuracy of symbol resolution. We implemented our method as part of the CodeCompass open source code comprehension tool and measured its efficiency.
Richárd Szalay, Zoltán Porkoláb, Dániel Krupp
SCAM1