Norman Peitek

dblp:120/4147 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-7828-4558ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Eye-Tracking Insights into the Effects of Type Annotations and Identifier Naming
Nils Alznauer, Norman Peitek, Youssef Abdelsalam, Annabelle Bergum, Marvin Wyrich, Sven Apel
ICPC2
2026 The Effect of Comments on Program Comprehension: An Eye-tracking Study
abstract
Abstract Programmers rely on code documentation and comments to understand source code, with program comprehension tasks consuming a significant portion of development time. Despite their importance, the impact of comments on program comprehension remains debated. Our study addresses this gap by investigating the influence of comments on program comprehension. Employing a mixed-methods approach, we conducted an eye-tracking study involving 20 computer science students to explore the influence of code comments on program comprehension. By analyzing both quantitative and qualitative data, we aim at comprehensively assessing the influence of comments on various aspects of program comprehension. The quantitative data collected consists of behavioral metrics assessing program comprehension in terms of correctness and response time, along with gaze data providing insights into visual attention, linearity of reading order, and gaze strategies. This was complemented by the participants’ ratings on the perceived difficulty and contribution of comments. Additionally, we gathered participants’ experiences through a post-questionnaire, enriching the analysis with qualitative insights into the effectiveness of comments, navigation strategies, and overall experiences with comments. Our findings reveal that the effect of comments on supporting program comprehension varies significantly across code snippets, ranging from a 30% decrease to a 34% increase in performance. Comments significantly guide visual attention, accounting for up to 23% of fixations, and promoted a more linear reading approach. Participants predominantly adhered to a “code-first” strategy. Moreover, comments were rated positively for clarifying complex segments of code and contributing to program comprehension. However, this favorable perception did not consistently translate into improved performance or reduced perceived difficulty across snippets. Based on our findings, we propose avenues for future research, including comparative studies on automated versus human-generated comments and the development of predictive models for assessing comment usefulness.
Youssef Abdelsalam, Norman Peitek, Annabelle Bergum, Sven Apel
Empir. Softw. Eng.2
2026 On the Influence of the Baseline in Neuroimaging Experiments on Program Comprehension
abstract
Background : Neuroimaging methods have been proved insightful in program-comprehension research. A key problem is that different baselines have been used in different experiments. A baseline is a task during which the “normal” brain activation is captured as a reference compared to the task of interest. Unfortunately, the influence of the choice of the baseline is still unclear. Aims : We investigate whether and to what extent the selected baseline influences the results of neuroimaging experiments on program comprehension. This helps to understand the tradeoffs in baseline selection with the ultimate goal of making the baseline selection informed and transparent. Method : We have conducted a pre-registered program-comprehension study with 20 participants using multiple baselines (i.e., reading, calculations, problem solving, and cross-fixation). We monitored brain activation with a 64-channel electroencephalography (EEG) device. We compared how the different baselines affect the results regarding brain activation of program comprehension. Results and Implications : We found significant differences in mental load across baselines suggesting that selecting a suitable baseline is critical. Our results show that a standard problem-solving task, operationalized by the Raven-Progressive Matrices, is a well-suited default baseline for program-comprehension studies. Our results highlight the need for carefully designing and selecting a baseline in program-comprehension studies.
Annabelle Bergum, Norman Peitek, Maurice Rekrut, Janet Siegmund, Sven Apel
ACM Trans. Softw. Eng. Methodol.2
2024 Data Analysis Tools Affect Outcomes of Eye-Tracking Studies
abstract
Background: Eye-tracking studies in software engineering offer insights into the thought processes of developers and their interaction with visual information. However, the analysis of eye-tracking data lacks established standards. Particularly concerning for the comparability of study findings is the large variety of tools for evaluating eye-tracking data, all of which use different algorithms and preset parameters, often undisclosed.
Timon Dörzapf, Norman Peitek, Marvin Wyrich, Sven Apel
ESEM2
2022 Correlates of programmer efficacy and their link to experience: a combined EEG and eye-tracking study
abstract
Background: Despite similar education and background, programmers can exhibit vast differences in efficacy. While research has identified some potential factors, such as programming experience and domain knowledge, the effect of these factors on programmers' efficacy is not well understood. Aims: We aim at unraveling the relationship between efficacy (speed and correctness) and measures of programming experience. We further investigate the correlates of programmer efficacy in terms of reading behavior and cognitive load. Method: For this purpose, we conducted a controlled experiment with 37 participants using electroencephalography (EEG) and eye tracking. We asked participants to comprehend up to 32 Java source-code snippets and observed their eye gaze and neural correlates of cognitive load. We analyzed the correlation of participants' efficacy with popular programming experience measures. Results: We found that programmers with high efficacy read source code more targeted and with lower cognitive load. Commonly used experience levels do not predict programmer efficacy well, but self-estimation and indicators of learning eagerness are fairly accurate. Implications: The identified correlates of programmer efficacy can be used for future research and practice (e.g., hiring). Future research should also consider efficacy as a group sampling method, rather than using simple experience measures.
Norman Peitek, Annabelle Bergum, Maurice Rekrut, Jonas Mucke, Matthias Nadig, Chris Parnin, Janet Siegmund, Sven Apel
ESEC/SIGSOFT FSE1
2021 Program Comprehension and Code Complexity Metrics: An fMRI Study
abstract
Background: Researchers and practitioners have been using code complexity metrics for decades to predict how developers comprehend a program. While it is plausible and tempting to use code metrics for this purpose, their validity is debated, since they rely on simple code properties and rarely consider particularities of human cognition. Aims: We investigate whether and how code complexity metrics reflect difficulty of program comprehension. Method: We have conducted a functional magnetic resonance imaging (fMRI) study with 19 participants observing program comprehension of short code snippets at varying complexity levels. We dissected four classes of code complexity metrics and their relationship to neuronal, behavioral, and subjective correlates of program comprehension, overall analyzing more than 41 metrics. Results: While our data corroborate that complexity metrics can-to a limited degree-explain programmers' cognition in program comprehension, fMRI allowed us to gain insights into why some code properties are difficult to process. In particular, a code's textual size drives programmers' attention, and vocabulary size burdens programmers' working memory. Conclusion: Our results provide neuro-scientific evidence supporting warnings of prior research questioning the validity of code complexity metrics and pin down factors relevant to program comprehension. Future Work: We outline several follow-up experiments investigating fine-grained effects of code complexity and describe possible refinements to code complexity metrics.
Norman Peitek, Sven Apel, Chris Parnin, André Brechmann, Janet Siegmund
ICSE1
2021 Mastering Variation in Human Studies: The Role of Aggregation
abstract
The human factor is prevalent in empirical software engineering research. However, human studies often do not use the full potential of analysis methods by combining analysis of individual tasks and participants with an analysis that aggregates results over tasks and/or participants. This may hide interesting insights of tasks and participants and may lead to false conclusions by overrating or underrating single-task or participant performance. We show that studying multiple levels of aggregation of individual tasks and participants allows researchers to have both insights from individual variations as well as generalized, reliable conclusions based on aggregated data. Our literature survey revealed that most human studies perform either a fully aggregated analysis or an analysis of individual tasks. To show that there is important, non-trivial variation when including human participants, we reanalyze 12 published empirical studies, thereby changing the conclusions or making them more nuanced. Moreover, we demonstrate the effects of different aggregation levels by answering a novel research question on published sets of fMRI data. We show that when more data are aggregated, the results become more accurate. This proposed technique can help researchers to find a sweet spot in the tradeoff between cost of a study and reliability of conclusions.
Janet Siegmund, Norman Peitek, Sven Apel, Norbert Siegmund
ACM Trans. Softw. Eng. Methodol.2
2020 What Drives the Reading Order of Programmers?: An Eye Tracking Study
abstract
Background: The way how programmers comprehend source code depends on several factors, including the source code itself and the programmer. Recent studies showed that novice programmers tend to read source code more like natural language text, whereas experts tend to follow the program execution flow. But, it is unknown how the linearity of source code and the comprehension strategy influence programmers' linearity of reading order.
Norman Peitek, Janet Siegmund, Sven Apel
ICPC1
2020 A Look into Programmers' Heads
abstract
Program comprehension is an important, but hard to measure cognitive process. This makes it difficult to provide suitable programming languages, tools, or coding conventions to support developers in their everyday work. Here, we explore whether functional magnetic resonance imaging (fMRI) is feasible for soundly measuring program comprehension. To this end, we observed 17 participants inside an fMRI scanner while they were comprehending source code. The results show a clear, distinct activation of five brain regions, which are related to working memory, attention, and language processing, which all fit well to our understanding of program comprehension. Furthermore, we found reduced activity in the default mode network, indicating the cognitive effort necessary for program comprehension. We also observed that familiarity with Java as underlying programming language reduced cognitive effort during program comprehension. To gain confidence in the results and the method, we replicated the study with 11 new participants and largely confirmed our findings. Our results encourage us and, hopefully, others to use fMRI to observe programmers and, in the long run, answer questions, such as: How should we train programmers? Can we train someone to become an excellent programmer? How effective are new languages and tools for program comprehension?
Norman Peitek, Janet Siegmund, Sven Apel, Christian Kästner, Chris Parnin, Anja Bethmann, Thomas Leich, Gunter Saake, André Brechmann
IEEE Trans. Software Eng.1
2019 Indentation: simply a matter of style or support for program comprehension?
abstract
An early study showed that indentation is not a matter of style, but provides actual support for program comprehension. In this paper, we present a non-exact replication of this study. Our aim is to provide empirical evidence for the suggested level of indentation made by many style guides. Following Miara and others, we also included the perceived difficulty, and we extended the original design to gain additional insights into the influence of indentation on visual effort by employing an eye-tracker. In the course of our study, we asked 22 participants to calculate the output of Java code snippets with different levels of indentation, while we recorded their gaze behavior. We did not find any indication that the indentation levels affect program comprehension or visual effort, so we could not replicate the findings of Miara and others. Nevertheless, our modernization of the original experiment design is a promising starting point for future studies in this field.
Jennifer Bauer, Janet Siegmund, Norman Peitek, Johannes C. Hofmeister, Sven Apel
ICPC3
2019 CodersMUSE: multi-modal data exploration of program-comprehension experiments
abstract
Program comprehension is a central cognitive process in programming. It has been in the focus of researchers for decades, but is still not thoroughly unraveled. Multi-modal psycho-physiological and neurobiological measurement methods have proved successful to gain a more holistic understanding of program comprehension. However, there is no proper tool support that lets researchers explore synchronized, conjoint multi-modal data, specifically designed for the needs in program-comprehension research. In this paper, we present CodersMUSE, a prototype implementation that aims to satisfy this crucial need.
Norman Peitek, Sven Apel, André Brechmann, Chris Parnin, Janet Siegmund
ICPC1
2018 Simultaneous measurement of program comprehension with fMRI and eye tracking: a case study
abstract
Background Researchers have recently started to validate decades-old program-comprehension models using functional magnetic resonance imaging (fMRI). While fMRI helps us to understand neural correlates of cognitive processes during program comprehension, its comparatively low temporal resolution (i.e., seconds) cannot capture fast cognitive subprocesses (i.e., milliseconds).
Norman Peitek, Janet Siegmund, Chris Parnin, Sven Apel, Johannes C. Hofmeister, André Brechmann
ESEM1
2017 Measuring neural efficiency of program comprehension
abstract
Most modern software programs cannot be understood in their entirety by a single programmer. Instead, programmers must rely on a set of cognitive processes that aid in seeking, filtering, and shaping relevant information for a given programming task. Several theories have been proposed to explain these processes, such as ``beacons,' for locating relevant code, and ``plans,'' for encoding cognitive models. However, these theories are decades old and lack validation with modern cognitive-neuroscience methods. In this paper, we report on a study using functional magnetic resonance imaging (fMRI) with 11 participants who performed program comprehension tasks. We manipulated experimental conditions related to beacons and layout to isolate specific cognitive processes related to bottom-up comprehension and comprehension based on semantic cues. We found evidence of semantic chunking during bottom-up comprehension and lower activation of brain areas during comprehension based on semantic cues, confirming that beacons ease comprehension.
Janet Siegmund, Norman Peitek, Chris Parnin, Sven Apel, Johannes C. Hofmeister, Christian Kästner, Andrew Begel, Anja Bethmann, André Brechmann
ESEC/SIGSOFT FSE2