EDBT 2026 Demo / reviewers in the wild / expert
Benjamin Hummel
dblp:78/295
· DBLP profile ↗
21ranked-venue papers
4as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 3 first-author · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
8 papers |
Empirical software engineering · 47% Software maintenance and evolution · 29% Requirements engineering and software design · 12% | |
| Network and information security
1 paper |
Blockchain and cryptocurrency security · 50% Systems and software security · 50% | |
| Computer networks
1 paper |
Routing and switching · 87% Internet architecture and protocols · 13% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Systems and software security › vulnerability discovery › machine-learning-based vulnerability detection
LLM-based vulnerability detection |
0.9 | 1 | 2025 | Should We Evaluate LLM Based Security Analysis Approaches on Open Source Systems? · ASE 2025 |
Blockchain and cryptocurrency security › smart contract security
vulnerability detection |
0.9 | 1 | 2025 | Should We Evaluate LLM Based Security Analysis Approaches on Open Source Systems? · ASE 2025 |
Software maintenance and evolution
code clone detection |
0.5 | 5 | 2010 | Can clone detection support quality assessments of requirements specifications? · ICSE (2) 2010 Code clone detection in practice · ICSE (2) 2010 Do code clones matter? · ICSE 2009 |
Requirements engineering and software design › software architecture › software architecture evolution
architectural decay |
0.1 | 1 | 2010 | Flexible architecture conformance assessment with ConQAT · ICSE (2) 2010 |
Program analysis › code quality analysis
redundancy detection |
0.1 | 1 | 2010 | Can clone detection support quality assessments of requirements specifications? · ICSE (2) 2010 |
Requirements engineering and software design
software architecture |
0.1 | 1 | 2010 | Flexible architecture conformance assessment with ConQAT · ICSE (2) 2010 |
Software testing
model-based testing |
0.1 | 1 | 2009 | Specifying the worst case: orthogonal modeling of hardware errors · ISSTA 2009 |
Software maintenance and evolution › code clone detection
model clone detection |
0.1 | 1 | 2008 | Clone detection in automotive model-based development · ICSE 2008 |
Routing and switching › inter-domain routing
AS relationship inference |
0.1 | 1 | 2007 | Acyclic type-of-relationship problems on the internet: an experimental analysis · Internet Measurement Conference 2007 |
Routing and switching
inter-domain routing |
0.1 | 1 | 2007 | Acyclic type-of-relationship problems on the internet: an experimental analysis · Internet Measurement Conference 2007 |
Compilers and program optimization
code duplication |
0.0 | 1 | 2010 | Code clone detection in practice · ICSE (2) 2010 |
Empirical software engineering
mining software repositories |
0.0 | 1 | 2009 | Do code clones matter? · ICSE 2009 |
Requirements engineering and software design
model-driven engineering |
0.0 | 1 | 2008 | Clone detection in automotive model-based development · ICSE 2008 |
Embedded and real-time systems
model-based design |
0.0 | 1 | 2008 | Clone detection in automotive model-based development · ICSE 2008 |
Internet architecture and protocols › network topology
autonomous system topology |
0.0 | 1 | 2007 | Acyclic type-of-relationship problems on the internet: an experimental analysis · Internet Measurement Conference 2007 |
Methods — techniques the papers use, named apart from their topics
large language model evaluation · 1.7fault injection · 0.2behavior modeling · 0.2graph theory · 0.2clone detection · 0.1heuristic · 0.1approximation algorithm · 0.1NP-hardness analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Should We Evaluate LLM Based Security Analysis Approaches on Open Source Systems?abstractExisting research has demonstrated promising results when applying large language models (LLMs) to detect security vulnerabilities in source code. However, these studies have been exclusively evaluated on benchmarks from open-source systems, using publicly known vulnerabilities that are likely part of the LLMs’ training data. This raises concerns that reported performance metrics may be inflated due to data contamination, providing a misleading view of the models’ actual capabilities.In this paper, we quantify this effect with a case study that evaluates five frontier LLMs on two carefully curated datasets: CWE-Bench-Java (an open-source dataset) and TS-Vuls (based on a closed-source commercial codebase). To provide a second angle, we also split CWE-Bench-Java by CVE record date to explore temporal contamination based on LLM knowledge cutoff dates.Our results reveal that the average F1 score dropped by approximately 20 percentage points when comparing the open-source to the closed-source dataset. Additionally, the precision drops from 56% to 34% on average, which is statistically significant (p < 0.05) for four of five models. This declining trend is consistent across all tested LLMs and metrics. In contrast, the results for the temporal split on the open-source data are inconclusive, suggesting that using a knowledge cutoff may reduce but does not ensure the elimination of contamination effects.Although our study is based on a single closed-source system and thus not generalizable, these findings provide the first empirical evidence that evaluating LLM-based vulnerability detection on open-source benchmarks may lead to overly optimistic results. This motivates the inclusion of closed-source datasets in future LLM evaluations. Kohei Dozono, Jonas Engesser, Benjamin Hummel, Tobias Roehm, Alexander Pretschner |
ASE | 3 |
| 2014 | Incremental origin analysis of source code filesabstractThe history of software systems tracked by version control systems is often incomplete because many file movements are not recorded. However, static code analyses that mine the file history, such as change frequency or code churn, produce precise results only if the complete history of a source code file is available. In this paper, we show that up to 38.9% of the files in open source systems have an incomplete history, and we propose an incremental, commit-based approach to reconstruct the history based on clone information and name similarity. With this approach, the history of a file can be reconstructed across repository boundaries and thus provides accurate information for any source code analysis. We evaluate the approach in terms of correctness, completeness, performance, and relevance with a case study among seven open source systems and a developer survey. Daniela Steidl, Benjamin Hummel, Elmar Jürgens |
MSR | 2 |
| 2013 | Quality analysis of source code commentsabstractA significant amount of source code in software systems consists of comments, i. e., parts of the code which are ignored by the compiler. Comments in code represent a main source for system documentation and are hence key for source code understanding with respect to development and maintenance. Although many software developers consider comments to be crucial for program understanding, existing approaches for software quality analysis ignore system commenting or make only quantitative claims. Hence, current quality analyzes do not take a significant part of the software into account. In this work, we present a first detailed approach for quality analysis and assessment of code comments. The approach provides a model for comment quality which is based on different comment categories. To categorize comments, we use machine learning on Java and C/C++ programs. The model comprises different quality aspects: by providing metrics tailored to suit specific categories, we show how quality aspects of the model can be assessed. The validity of the metrics is evaluated with a survey among 16 experienced software developers, a case study demonstrates the relevance of the metrics in practice. Daniela Steidl, Benjamin Hummel, Elmar Jürgens |
ICPC | 2 |
| 2013 | Behavioral specification of reactive systems using stream-based I/O tables
Judith Thyssen, Benjamin Hummel |
Softw. Syst. Model. | 2 |
| 2012 | A framework for incremental quality analysis of large software systemsabstractTo provide rapid feedback to engineers, software quality analysis must be incremental. However, most existing analyses are either not incremental, or limited to isolated quality characteristics. In practice, this prevents their integration into a uniform quality control approach. In this paper, we present a framework for the incremental and distributed computation of quality characteristics. It is fast enough for real-time analysis of large systems and provides a complete history of analysis results. An evaluation on several open source software systems demonstrates its scalability to large code bases under active development. Veronika Bauer, Lars Heinemann, Benjamin Hummel, Elmar Jürgens, Michael Conradt |
ICSM | 3 |
| 2011 | On the Extent and Nature of Software Reuse in Open Source Java Projects
Lars Heinemann, Florian Deißenböck, Mario Gleirscher, Benjamin Hummel, Maximilian Irlbeck |
ICSR | 4 |
| 2011 | Semantic Clone Detection for Model-Based Development of Embedded Systems
Bakr Al-Batran, Bernhard Schätz, Benjamin Hummel |
MoDELS | 3 |
| 2010 | Flexible architecture conformance assessment with ConQATabstractThe architecture of software systems is known to decay if no counter-measures are taken. In order to prevent this architectural erosion, the conformance of the actual system architecture to its intended architecture needs to be assessed and controlled; ideally in a continuous manner. To support this, we present the architecture conformance assessment capabilities of our quality analysis framework ConQAT. In contrast to other tools, ConQAT is not limited to the assessment of use-dependencies between software components. Its generic architectural model allows the assessment of various types of dependencies found between different kinds of artifacts. It thereby provides the necessary tool-support for flexible architecture conformance assessment in diverse contexts. Florian Deißenböck, Lars Heinemann, Benjamin Hummel, Elmar Jürgens |
ICSE (2) | 3 |
| 2010 | Code clone detection in practiceabstractDue to the negative impact of code cloning on software maintenance efforts as well as on program correctness [4--6], the duplication of code is generally viewed as problematic. However, the techniques and tools developed by the research community in the last decade have not found broad acceptance in software engineering practice yet. This tutorial contributes to a more widespread application of existing approaches by illustrating where cloning comes from, what its consequences are, and how it can be detected. Florian Deißenböck, Benjamin Hummel, Elmar Jürgens |
ICSE (2) | 2 |
| 2010 | Can clone detection support quality assessments of requirements specifications?abstractDue to their pivotal role in software engineering, considerable effort is spent on the quality assurance of software requirements specifications. As they are mainly described in natural language, relatively few means of automated quality assessment exist. However, we found that clone detection, a technique widely applied to source code, is promising to assess one important quality aspect in an automated way, namely redundancy that stems from copy&paste operations. This paper describes a large-scale case study that applied clone detection to 28 requirements specifications with a total of 8,667 pages. We report on the amount of redundancy found in real-world specifications, discuss its nature as well as its consequences and evaluate in how far existing code clone detection approaches can be applied to assess the quality of requirements specifications in practice. Elmar Jürgens, Florian Deißenböck, Martin Feilkas, Benjamin Hummel, Bernhard Schätz, Stefan Wagner 0001, Christoph Domann, Jonathan Streit |
ICSE (2) | 4 |
| 2010 | Index-based code clone detection: incremental, distributed, scalableabstractAlthough numerous different clone detection approaches have been proposed to date, not a single one is both incremental and scalable to very large code bases. They thus cannot provide real-time cloning information for clone management of very large systems. We present a novel, index-based clone detection algorithm for type 1 and 2 clones that is both incremental and scalable. It enables a new generation of clone management tools that provide real-time cloning information for very large software. We report on several case studies that show both its suitability for real-time clone detection and its scalability: on 42 MLOC of Eclipse code, average time to retrieve all clones for a file was below 1 second; on 100 machines, detection of all clones in 73 MLOC was completed in 36 minutes. Benjamin Hummel, Elmar Jürgens, Lars Heinemann, Michael Conradt |
ICSM | 1 |
| 2010 | Material Flow Abstraction of Manufacturing Systems
Jewgenij Botaschanjan, Benjamin Hummel |
ICTAC | 2 |
| 2010 | A Generalized Wedelin Heuristic for Integer ProgrammingabstractA very important ingredient for solving hard general integer programs are heuristics that try to quickly find good feasible solutions. One of these heuristics is Wedelin's algorithm, which works for the limited class of 0-1 integer programs. A big advantage of Wedelin's approach is that it does not depend on a solution of the linear programming (LP) relaxation as many other heuristics do. This makes it extremely fast in practice and makes it easy to use the parallelism of the upcoming multicore CPUs, as in an integer programming (IP) solver it could be applied in parallel to the traditional branch-and-bound algorithm. In this paper, we present several extensions and generalizations to Wedelin's algorithm (most can be handled in an implicit manner without much performance cost) and investigate different ways of improving it. We give all necessary details and parameters. We strive for an algorithm that is faster than other heuristics but achieves comparable solution quality. We evaluate the performance of the algorithm on a large set of more than 100 instances from different sources. The results indicate that our heuristic often finds solutions comparable to or even better than those found using current state-of-the-art heuristics while typically needing only a fraction of their running time. Additionally, we report positive findings on the application of the heuristic on feasibility instances from discrete tomography. Our algorithm always finds the IP optimum in less time than the simplex/barrier algorithms and often in less time than it takes the volume algorithm to find just the LP optimum. Oliver Bastert, Benjamin Hummel, Sven de Vries |
INFORMS J. Comput. | 2 |
| 2009 | Integrated Behavior Models for Factory Automation SystemsabstractDespite the large amount of models for different aspects of factory automation systems, many of these models target at individual and in most cases static aspects of the system, such as the geometry or its electric parts. There is a lack of suitable description methods, which integrate these individual models to a behavior model including spatial aspects and the handling of material. Furthermore, it is important that this model keeps the link to the more detailed individual models and is sufficiently formal in order to allow an automated analysis. This paper provides a solution to this problem by introducing a model which addresses both spatial structure and behavior and is based on a thorough mathematical theory. Complementary, we report on a tool realization of the modelling theory and explain how the model supports the development of mechatronic systems. Jewgenij Botaschanjan, Benjamin Hummel, Thomas Hensel, Alexander Lindworsky |
ETFA | 2 |
| 2009 | CloneDetective - A workbench for clone detection researchabstractThe area of clone detection has considerably evolved over the last decade, leading to approaches with better results, but at the same time using more elaborate algorithms and tool chains. In our opinion a level has been reached, where the initial investment required to setup a clone detection tool chain and the code infrastructure required for experimenting with new heuristics and algorithms seriously hampers the exploration of novel solutions or specific case studies. As a solution, this paper presents CloneDetective, an open source framework and tool chain for clone detection, which is especially geared towards configurability and extendability and thus supports the preparation and conduction of clone detection research. Elmar Jürgens, Florian Deißenböck, Benjamin Hummel |
ICSE | 3 |
| 2009 | Do code clones matter?abstractCode cloning is not only assumed to inflate maintenance costs but also considered defect-prone as inconsistent changes to code duplicates can lead to unexpected behavior. Consequently, the identification of duplicated code, clone detection, has been a very active area of research in recent years. Up to now, however, no substantial investigation of the consequences of code cloning on program correctness has been carried out. To remedy this shortcoming, this paper presents the results of a large-scale case study that was undertaken to find out if inconsistent changes to cloned code can indicate faults. For the analyzed commercial and open source systems we not only found that inconsistent changes to clones are very frequent but also identified a significant number of faults induced by such changes. The clone detection tool used in the case study implements a novel algorithm for the detection of inconsistent clones. It is available as open source to enable other researchers to use it as basis for further investigations. Elmar Jürgens, Florian Deißenböck, Benjamin Hummel, Stefan Wagner 0001 |
ICSE | 3 |
| 2009 | Specifying the worst case: orthogonal modeling of hardware errorsabstractDuring testing, the execution of valid cases is only one part of the task. Checking the behavior in boundary situations and in the presence of errors is an equally important subject. This is especially true in embedded systems where parts of a system's function are realized by sensors and actuators, which are subject to wear and defects. As testing with the real hardware is costly and hardware defects are hard to stimulate, such tests are often performed using behavior models of the system which allow to execute the controller software against simulated hardware and environment. However, these models seldom contain possible hardware errors, as this makes the models more complex and, thus, harder to create and maintain. This paper presents a modeling technique for the description of system errors without modifying the original model. Error specifications for individual system components are modeled separately and can be used to augment the system model. Jewgenij Botaschanjan, Benjamin Hummel |
ISSTA | 2 |
| 2009 | Behavioral Specification of Reactive Systems Using Stream-Based I/O TablesabstractA core problem in formal methods is the transition from informal requirements to formal specifications. Especially when specifying reactive systems, many formalisms require the user to either understand a complex mathematical theory and notation or to derive details not given in the requirements, such as the state space of the problem. While formalizing a real-world requirements document, we developed a technique where not states but signal patterns are the main elements. We argue that it supports a formalization that is often closer to the informal requirements and thus provides a smoother transition to formal methods. As only tables of regular expressions are used for notation, the technique can easily be understood by non-mathematicians. Many properties, such as consistency, can be checked automatically on these specifications. Besides the formal foundation of our approach, this paper presents prototypical tool support and first results from an industrial case study. Benjamin Hummel, Judith Thyssen |
SEFM | 1 |
| 2008 | Clone detection in automotive model-based developmentabstractModel-based development is becoming an increasingly common development methodology. In important domains like embedded systems already major parts of the code are generated from models specified with domain-specific modelling languages. Hence, such models are nowadays an integral part of the software development and maintenance process and therefore have a major economic and strategic value for the software-developing organisations. Nevertheless almost no work has been done on a quality defect that is known to seriously hamper maintenance productivity in classic code-based development: Cloning. This paper presents an approach for the automatic detection of clones in large models as they are used in model-based development of control systems. The approach is based on graph theory and hence can be applied to most graphical data-flow languages. An industrial case study demonstrates the applicability of our approach for the detection of clones in Matlab/Simulink models that are widely used in model-based development of embedded systems in the automotive domain. Florian Deißenböck, Benjamin Hummel, Elmar Jürgens, Bernhard Schätz, Stefan Wagner 0001, Jean-Francois Girard, Stefan Teuchert |
ICSE | 2 |
| 2008 | Towards an integrated system model for testing and verification of automation machinesabstractModels and documents created during the development of automation machines typically can be categorized into mechanics, electronics, and software. The functionality of an automation machine is, however, realized by the interaction of all three of these domains. So no single model covering only one development category will be able to describe the behavior of the machine thoroughly. For early planning of machine design, virtual prototypes, and especially for the formal verification of requirements an integrated functional model of the machine is required. This paper introduces a technique which can be used to model automation machines on an abstract level, including coarse-grained descriptions of mechanics, electronics and software aspects with special focus on modeling domain-specific issues such as material flow and collision response. The resulting models are detailed enough to be simulated or verified but still suitably abstract to allow fast creation and efficient simulation. Benjamin Hummel, Peter Braun 0003 |
MiSE | 1 |
| 2007 | Acyclic type-of-relationship problems on the internet: an experimental analysisabstractAn experimental study of the feasibility and accuracy of the acyclicity approach introduced in [14] for the inference of business relationships among autonomous systems (ASes) is provided. We investigate the maximum acyclic type-of-relationship problem: on a given set of AS paths, find a maximum-cardinality subset which allows an acyclic and valley-free orientation. Inapproximability and NP-hardness results for this problem are presented and a heuristic is designed. The heuristic is experimentally compared to most of the state-of-the-art algorithms on a reliable data set. It turns out that the proposed heuristic produces the least number of misclassified customer-to-provider relationships among the tested algorithms. Moreover, it is flexible in handling pre-knowledge in the sense that already a small amount of correct relationships is enough to produce a high-quality relationship classification. Furthermore, the reliable data set is used to validate the acyclicity assumptions. The findings demonstrate that acyclicity notions should be an integral part of models of AS relationships. Benjamin Hummel, Sven Kosub |
Internet Measurement Conference | 1 |