Robert Heumüller

dblp:158/8506 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-9906-0323ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Empirical Analysis of OpenAI Embeddings for Semantic Code Review Comment Similarity
Robert Heumüller, Theo Langer, Frank Ortmeier
SEAA1
2021 Exploit those code reviews! bigger data for deeper learning
abstract
Modern code review (MCR) processes are prevalent in most organizations that develop software due to benefits in quality assurance and knowledge transfer. With the rise of collaborative software development platforms like GitHub and Bitbucket, today, millions of projects share not only their code but also their review data. Although researchers have tried to exploit this data for more than a decade, most of that knowledge remains a buried treasure. A crucial catalyst for many advances in deep learning, however, is the accessibility of large-scale standard datasets for different learning tasks. This paper presents the ETCR (Exploit Those Code Reviews!) infrastructure for mining MCR datasets from any GitHub project practicing pull-request-based development. We demonstrate its effectiveness with ETCR-Elasticsearch, a dataset of >231𝑘 review comments for >47𝑘 Java file revisions in >40𝑘 pull-requests from the Elasticsearch project. ETCR is designed with the challenge of deep learning in mind. Compared to previous datasets, ETCR datasets include all information for linking review comments to nodes in the respective program’s Abstract Syntax Tree.
Robert Heumüller, Sebastian Nielebock 0001, Frank Ortmeier
ESEC/SIGSOFT FSE1
2021 Guided pattern mining for API misuse detection by change-based code analysis
abstract
Abstract Lack of experience, inadequate documentation, and sub-optimal API design frequently cause developers to make mistakes when re-using third-party implementations. Such API misuses can result in unintended behavior, performance losses, or software crashes. Therefore, current research aims to automatically detect such misuses by comparing the way a developer used an API to previously inferred patterns of the correct API usage. While research has made significant progress, these techniques have not yet been adopted in practice. In part, this is due to the lack of a process capable of seamlessly integrating with software development processes. Particularly, existing approaches do not consider how to collect relevant source code samples from which to infer patterns. In fact, an inadequate collection can cause API usage pattern miners to infer irrelevant patterns which leads to false alarms instead of finding true API misuses. In this paper, we target this problem (a) by providing a method that increases the likelihood of finding relevant and true-positive patterns concerning a given set of code changes and agnostic to a concrete static, intra-procedural mining technique and (b) by introducing a concept for just-in-time API misuse detection which analyzes changes at the time of commit. Particularly, we introduce different, lightweight code search and filtering strategies and evaluate them on two real-world API misuse datasets to determine their usefulness in finding relevant intra-procedural API usage patterns. Our main results are (1) commit-based search with subsequent filtering effectively decreases the amount of code to be analyzed, (2) in particular method-level filtering is superior to file-level filtering, (3) project-internal and project-external code search find solutions for different types of misuses and thus are complementary, (4) incorporating prior knowledge of the misused API into the search has a negligible effect.
Sebastian Nielebock 0001, Robert Heumüller, Kevin Michael Schott, Frank Ortmeier
Autom. Softw. Eng.2
2020 How Can I Contribute?: A Qualitative Analysis of Community Websites of 25 Unix-Like Distributions
abstract
Developers collaboratively implement large-scale industrial and open-source projects. Such projects pose several challenges for developers, as they require considerable knowledge about the project and its development processes, for instance, to fix bugs or implement new features. Understanding what information developer communities codify on how to contribute to their project is crucial, for example, to onboard new developers or for researchers to scope analysis techniques. In this paper, we report the results of a qualitative analysis of 25 Unix-like distributions, focusing on what information the communities codify publicly on contributing. The results reveal no dedicated strategies to codify information on contribution or development practices. Still, non-technical contributions are easy to identify, while information on the development is hard to collect---and mostly concerned with versioning and bug reporting. Our insights help to understand information-provisioning strategies, identify information sources, and scope analyses.
Jacob Krüger, Sebastian Nielebock 0001, Robert Heumüller
EASE3
2020 Publish or perish, but do not forget your software artifacts
abstract
Abstract Open-science initiatives have gained substantial momentum in computer science, and particularly in software-engineering research. A critical aspect of open-science is the public availability of artifacts (e.g., tools), which facilitates the replication, reproduction, extension, and verification of results. While we experienced that many artifacts are not publicly available, we are not aware of empirical evidence supporting this subjective claim. In this article, we report an empirical study on software artifact papers (SAPs) published at the International Conference on Software Engineering (ICSE), in which we investigated whether and how researchers have published their software artifacts, and whether this had scientific impact. Our dataset comprises 789 ICSE research track papers, including 604 SAPs (76.6 %), from the years 2007 to 2017. While showing a positive trend towards artifact availability, our results are still sobering. Even in 2017, only 58.5 % of the papers that stated to have developed a software artifact made that artifact publicly available. As we did find a small, but statistically significant, positive correlation between linking to artifacts in a paper and its scientific impact in terms of citations, we hope to motivate the research community to share more artifacts. With our insights, we aim to support the advancement of open science by discussing our results in the context of existing initiatives and guidelines. In particular, our findings advocate the need for clearly communicating artifacts and the use of non-commercial, persistent archives to provide replication packages.
Robert Heumüller, Sebastian Nielebock 0001, Jacob Krüger, Frank Ortmeier
Empir. Softw. Eng.1
2019 SpecTackle - A Specification Mining Experimentation Platform
abstract
Nowadays, API Specification Mining is an important cornerstone of automated software engineering. In this paper, we introduce SpecTackle, an IDE-based experimentation platform aiming to facilitate experimentation and validation of specification mining algorithms and tools. SpecTackle strives toward (1) providing easy access to various specification mining tools, (2) simplifying configuration and usage through a shared interface, and (3) in-code visualization of pattern occurrences. The first version supports two heterogeneous mining tools, a third-party graph-based miner as well as a custom sequence mining tool. In the long term, SpecTackle envisions to also provide ground-truth benchmark projects, a unified pattern meta-model and parameter optimization for mining tools.
Robert Heumüller, Sebastian Nielebock 0001, Frank Ortmeier
SEAA1
2019 Programmers do not favor lambda expressions for concurrent object-oriented code
Sebastian Nielebock 0001, Robert Heumüller, Frank Ortmeier
Empir. Softw. Eng.2