Matús Sulír

dblp:170/6263 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
3since 2021 · last 2026
0000-0003-2221-9225ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorArtificial intelligence and machine learning · 3 · 2 first-author
YearPublicationVenuePosition
2026 Local software buildability across Java versions
abstract
Abstract Context Downloading the source code of open-source Java projects and building them on a local computer using Maven, Gradle, or Ant is a common activity performed by researchers and practitioners. Several studies have so far found that about 40–60% of such attempts fail. Our experience in recent years suggests that the proportion of failed builds continually rises even further. Objective First, we empirically tested our hypothesis that with increasing Java versions, the percentage of build-failing projects tends to grow. Next, nine additional research questions were proposed, related mainly to the proportions of failing projects, universal version compatibility, failures under specific JDK versions, success rates of build tools, wrappers, and failure categories. Method We sampled 2,500 random plain Java projects having a build configuration file and meeting basic quality criteria from GitHub. We tried to automatically build every project in containers with Java versions 6 to 23 installed. Success or failure was determined by exit codes, and standard output and error streams were saved. Data were automatically processed and supplemented with manual analysis when necessary. Results Our hypothesis was not confirmed; there is no monotonic trend in the projects’ build success rates. From Java 8 onward, they range from 30 to 44%, with a peak at JDK 17. Downgrading the latest JDK can fix 52% of failed builds. Trying the long-term support versions (8, 11, 17, and 21) is a particularly effective build repair strategy. We also found that about 32% of projects fail for all JDKs. Gradle consistently has the lowest success rate of all tools, and builds most frequently fail during initialization and compilation. Conclusions Building projects with a wrong Java version is a significant failure factor. We provide actionable suggestions for language designers, build tool authors, project maintainers, researchers designing build repair tools, and developers compiling open-source projects.
Matús Sulír, Jaroslav Porubän, Sergej Chodarev
Empir. Softw. Eng.1
2023 Outside the Sandbox: A Study of Input/Output Methods in Java
abstract
Programming languages often demarcate the internal sandbox, consisting of entities such as objects and variables, from the outside world, e.g., files or network. Although communication with the external world poses fundamental challenges for live programming, reversible debugging, testing, and program analysis in general, studies about this phenomenon are rare. In this paper, we present a preliminary empirical study about the prevalence of input/output (I/O) method usage in Java. We manually categorized 1435 native methods in a Java Standard Edition distribution into non-I/O and I/O-related methods, which were further classified into areas such as desktop or file-related ones. According to the static analysis of a call graph for 798 projects, about 57% of methods potentially call I/O natives. The results of dynamic analysis on 16 benchmarks showed that 21% of the executed methods directly or indirectly called an I/O native. We conclude that neglecting I/O is not a viable option for tool designers and suggest the integration of I/O-related metadata with source code to facilitate their querying.
Matús Sulír, Sergej Chodarev, Milan Nosál
EASE1
2022 A fine-grained data set and analysis of tangling in bug fixing commits
abstract
Abstract Context Tangled commits are changes to software that address multiple concerns at once. For researchers interested in bugs, tangled commits mean that they actually study not only bugs, but also other concerns irrelevant for the study of bugs. Objective We want to improve our understanding of the prevalence of tangling and the types of changes that are tangled within bug fixing commits. Methods We use a crowd sourcing approach for manual labeling to validate which changes contribute to bug fixes for each line in bug fixing commits. Each line is labeled by four participants. If at least three participants agree on the same label, we have consensus. Results We estimate that between 17% and 32% of all changes in bug fixing commits modify the source code to fix the underlying problem. However, when we only consider changes to the production code files this ratio increases to 66% to 87%. We find that about 11% of lines are hard to label leading to active disagreements between participants. Due to confirmed tangling and the uncertainty in our data, we estimate that 3% to 47% of data is noisy without manual untangling, depending on the use case. Conclusion Tangled commits have a high prevalence in bug fixes and can lead to a large amount of noise in the data. Prior research indicates that this noise may alter results. As researchers, we should be skeptics and assume that unvalidated data is likely very noisy, until proven otherwise.
Steffen Herbold, Alexander Trautsch, Benjamin Ledel, Alireza Aghamohammadi, Taher Ahmed Ghaleb, Kuljit Kaur Chahal, Tim Bossenmaier, Bhaveet Nagaria, Philip Makedonski, Matin Nili Ahmadabadi, Kristóf Szabados, Helge Spieker, Matej Madeja, Nathaniel Hoy, Valentina Lenarduzzi, Shangwen Wang, Gema Rodríguez-Pérez, Ricardo Colomo-Palacios, Roberto Verdecchia, Paramvir Singh, Yihao Qin, Debasish Chakroborti, Willard Davis, Vijay Walunj, Diego Marcilio, Omar Alam, Abdullah Aldaeej, Idan Amit, Burak Turhan, Simon Eismann, Anna-Katharina Wickert, Ivano Malavolta, Matús Sulír, Fatemeh Hendijani Fard, Austin Z. Henley, Stratos Kourtzanidis, Eray Tüzün, Christoph Treude, Simin Maleki Shamasbi, Ivan Pashchenko, Marvin Wyrich, James C. Davis 0001, Alexander Serebrenik, Ella Albrecht, Ethem Utku Aktas, Daniel Strüber 0001, Johannes Erbel
Empir. Softw. Eng.34
2020 String Representations of Java Objects: An Empirical Study
Matús Sulír
SOFSEM1
2018 Integrating Runtime Values with Source Code to Facilitate Program Comprehension
abstract
An inherently abstract nature of source code makes programs difficult to understand. In our research, we designed three techniques utilizing concrete values of variables and other expressions during program execution. RuntimeSearch is a debugger extension searching for a given string in all expressions at runtime. DynamiDoc generates documentation sentences containing examples of arguments, return values and state changes. RuntimeSamp augments source code lines in the IDE (integrated development environment) with sample variable values. In this post-doctoral article, we briefly describe these three approaches and related motivational studies, surveys and evaluations. We also reflect on the PhD study, providing advice for current students. Finally, short-term and long-term future work is described.
Matús Sulír
ICSME1
2018 Augmenting source code lines with sample variable values
abstract
Source code is inherently abstract, which makes it difficult to understand. Activities such as debugging can reveal concrete runtime details, including the values of variables. However, they require that a developer explicitly requests these data for a specific execution moment. We present a simple approach, RuntimeSamp, which collects sample variable values during normal executions of a program by a programmer. These values are then displayed in an ambient way at the end of each line in the source code editor. We discuss questions which should be answered for this approach to be usable in practice, such as how to efficiently record the values and when to display them. We provide partial answers to these questions and suggest future research directions.
Matús Sulír, Jaroslav Porubän
ICPC1
2017 Labeling Source Code with Metadata: A Survey and Taxonomy
abstract
Source code is a primary artifact where programmers are looking when they try to comprehend a program.However, to improve program comprehension efficiency, tools often associate parts of source code with metadata collected from static and dynamic analysis, communication artifacts and many other sources.In this article, we present a systematic mapping study of approaches and tools labeling source code elements with metadata and presenting them to developers in various forms.We selected 25 from more than 2,000 articles and categorized them.A taxonomy with four dimensions -source, target, presentation and persistence -was formed.Based on the survey results, we also identified interesting future research challenges.
Matús Sulír, Jaroslav Porubän
FedCSIS1
2017 RuntimeSearch: Ctrl+F for a running program
abstract
Developers often try to find occurrences of a certain term in a software system. Traditionally, a text search is limited to static source code files. In this paper, we introduce a simple approach, RuntimeSearch, where the given term is searched in the values of all string expressions in a running program. When a match is found, the program is paused and its runtime properties can be explored with a traditional debugger. The feasibility and usefulness of RuntimeSearch is demonstrated on a medium-sized Java project.
Matús Sulír, Jaroslav Porubän
ASE1
2017 Customizing host IDE for non-programming users of pure embedded DSLs: A case study
Milan Nosál, Jaroslav Porubän, Matús Sulír
Comput. Lang. Syst. Struct.3
2016 Recording concerns in source code using annotations
Matús Sulír, Milan Nosál, Jaroslav Porubän
Comput. Lang. Syst. Struct.1
2015 Source code annotations as formal languages
abstract
Attribute-oriented programming (source code annotations) is a program level marking technique that enables enrichment of program elements with custom metadata.In this paper we hypothesize that there is a correspondence between source code annotations and conventional formal languages in general.We analyze our observations about source code annotations from three aspects of language description: concrete syntax, abstract syntax, and semantics.The discussion provides evidence of the hypothesized correspondence and we use it as a basis for our definition of an annotation-based language (abbreviated: @L).However, the analysis also shows that compared to conventional formal languages, source code annotations have some specificities mainly connected to their binding to host program elements.The presented analysis contributes to the field of attributeoriented programming by discussing the relationship between annotations and conventional formal languages, and by surveying relational idioms in annotations' usage that can be inspirational for annotations' authors.
Milan Nosál, Matús Sulír, Ján Juhár
FedCSIS2
2015 Sharing developers' mental models through source code annotations
abstract
Context: Developers possess mental models containing information far beyond what is explicitly captured in the source code.Objectives: We investigate the possibility to use source code annotations to capture parts of the developers' mental models and later reuse them by other programmers during program comprehension and maintenance.Method: We performed two studies and a controlled experiment.Results: Developers' mental models overlap and thus can be shared.Possible use cases of shared annotations are hypotheses confirmation, feature location, obtaining new knowledge, finding relationships and maintenance notes.In the experiment, the presence of annotations reduced program comprehension and maintenance time by 34%.Conclusion: Annotations are a viable way to share programmers' thoughts.
Matús Sulír, Milan Nosál
FedCSIS1