VLDB 2026 Research / reviewers in the wild / expert
Benedikt Fein
dblp:324/0477
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-3798-845XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Challenges of deploying code embeddings: an industrial case study on method name generationabstractAbstract The recent hype around machine learning has fully captured software engineering research. Correspondingly, a variety of different ways to represent code as input to deep learning models have been proposed. These code embedding models are usually evaluated in terms of common metrics such as accuracy or bleu scores, on benchmark tasks such as predicting method names from their body. Although this evaluation approach is well established in research, it leaves open challenges for the deployment of these models in practice: First, comparing accuracy on standardised benchmark results conveniently avoids some of the challenges of actually running different prototype model implementations, which, however is necessary to apply the models in practice. Second, the models are usually trained and evaluated on abundantly available open-source training data, which may be very different from closed-source industrial code. Third, the deployment of machine learning models in an industrial environment does not only entail technical but also organisational challenges. Finally, while competitive accuracy or bleu scores may be indicative of relative model performance, they may not reflect to what extent the models are suitable for being used by developers. In this paper we describe our experience of evaluating and deploying state-of-research code embedding models in an industrial environment, and present lessons learned from our struggles with each of these questions. Benedikt Fein, Maximilian Jungwirth, Gordon Fraser 0001, Florian Kandlinger |
Autom. Softw. Eng. | 1 |
| 2025 | AsserT5: Test Assertion Generation Using a Fine-Tuned Code Language ModelabstractWriting good software tests can be challenging, therefore approaches that support developers are desirable. While generating complete tests automatically is such an approach commonly proposed in research, developers may already have specific test scenarios in mind and thus just require help in selecting the most suitable test assertions for these scenarios. This can be done using deep learning models to predict assertions for given test code. Prior research on assertion generation trained these models specifically for the task, raising the question how much the use of larger models pre-trained on code that have emerged since then can improve their performance. In particular, while abstracting identifiers has been shown to improve specifically trained models, it remains unclear whether this also generalises to models pre-trained on non-abstracted code. Finally, even though prior work demonstrated high accuracy it remains unclear how this translates into the effectiveness of the assertions at their intended application – finding faults. To shed light on these open questions, in this paper we propose AsserT5, a new model based on the pre-trained CodeT5 model, and use this to empirically study assertion generation. We find that the abstraction and the inclusion of the focal method are useful also for a fine-tuned pre-trained model, resulting in test assertions that match the ground truth assertions precisely in up to 59.5% of cases, more than twice as precise as prior models. However, evaluation on real bugs from the Defects4J dataset shows that out of 138 bugs detectable with assertions in real-world projects, AsserT5 was only able to suggest fault-finding assertions for 33, indicating the need for further improvements. Severin Primbs, Benedikt Fein, Gordon Fraser 0001 |
AST | 2 |
| 2025 | LitterBox+: An Extensible Framework for LLM-enhanced Scratch Static Code Analysis
Benedikt Fein, Florian Obermüller, Gordon Fraser 0001 |
ASE | 1 |
| 2023 | On the Applicability of Language Models to Block-Based ProgramsabstractBlock-based programming languages like Scratch are increasingly popular for programming education and end-user programming. Recent program analyses build on the insight that source code can be modelled using techniques from natural language processing. Many of the regularities of source code that support this approach are due to the syntactic overhead imposed by textual programming languages. This syntactic overhead, however, is precisely what block-based languages remove in order to simplify programming. Consequently, it is unclear how well this modelling approach performs on block-based programming languages. In this paper, we investigate the applicability of language models for the popular block-based programming language Scratch. We model Scratch programs using n-gram models, the most essential type of language model, and transformers, a popular deep learning model. Evaluation on the example tasks of code completion and bug finding confirm that blocks inhibit predictability, but the use of language models is nevertheless feasible. Our findings serve as foundation for improving tooling and analyses for block-based languages. Elisabeth Griebl, Benedikt Fein, Florian Obermüller, Gordon Fraser 0001, René Just |
ICSE | 2 |
| 2022 | An Evaluation of code2vec Embeddings for Scratch
Benedikt Fein, Isabella Graßl, Florian Beck, Gordon Fraser 0001 |
EDM | 1 |
| 2022 | CATNIP: An Automated Hint Generation Tool for ScratchabstractTaking the first steps when learning how to program can be hard. Block-based programming languages like Scratch lower this hurdle, but learners may nevertheless get stuck when trying to solve a specific task and need help. This can also challenge teachers when facing many raised hands at the same time in the classroom. Consequently, it is desirable for learners and teachers alike to have access to systems that automatically generate hints on which steps to take next in a programming assignment. In this paper we introduce Catnip, a tool that generates next step hints for the Scratch programming language based on a structural comparison between model solutions and the current student attempt. Catnip uses extensive postprocessing to improve the generated hints, and displays them directly inside the Scratch framework, suggesting where to add or reorder blocks while working on a programming task. Benedikt Fein, Florian Obermüller, Gordon Fraser 0001 |
ITiCSE (1) | 1 |