Marco Klaiber

dblp:302/7522 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
16since 2021 · last 2025
0009-0007-7070-3413ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Engaging Students in Scientific Writing: The STRaWBERRY Checklist Framework with LLM-based Paper Draft Assessment
abstract
Writing scientific papers is essential for advancing any given research field, yet undergraduate and graduate students often struggle with this task, facing challenges in clearly presenting their ideas and results. Valuable scientific contributions described in papers that are not well-structured and not easy to follow may not get published. This paper addresses these challenges by proposing a framework called STRaWBERRY, which provides a checklist-based guide to help evaluate individual components of paper drafts against essential quality criteria. Additionally, we propose the use of Large Language Models (LLMs) to automate the assessment based on these criteria. The LLM evaluation process encourages active learning (in the educational sense) and allows the drafts to be iteratively refined through feedback from the LLM. We evaluate the STRaWBERRY framework by its use in lectures and the corresponding outcome: STRaWBERRY has been successfully used in 10 university courses, leading to multiple student pub-lications. Furthermore, we evaluate the LLM-based approach by assessing its accuracy in evaluating a selection of sample papers, demonstrating its potential to supplement and enhance traditional proofreading cycles.
Andreas Theissler, Marco Klaiber, Felix Gerschner, Philip Ritzer
EDUCON2
2025 Optimizing Aircraft Assembly: A Machine Learning Approach for Screw Fastener Outcome Prediction
abstract
Screw fastening is one of the most important manufacturing processes in the aerospace industry. For example, millions of fasteners are utilized in the Boeing 777 to connect around 3 million different structural components. However, screw fasteners mounted by articulated arm robots in the assembly of aircraft parts are subject to positional uncertainties in both dynamic and static applications. The resulting production errors not only have an impact on cost and time but also lead to safety risks. Despite this strong theoretical and practical relevance, only few research approaches focus on the recognition of defective fasteners in aircraft manufacturing, and the detection of incorrectly assembled aeronautical threaded fasteners remains open. To address these challenges, we identified important time series channels and developed a hard voting ensemble of Categorical Boosting (CatBoost) and Extreme Gradient Boosting (XGBoost) to predict the outcomes of screw fasteners mounted by articulated arm robots in the assembly of aircraft parts. Our model achieved a balanced accuracy of 81.08% on a time series dataset containing both kinematic and dynamic data obtained from a robotic assembly process of aeronautical screwed fasteners. While both CatBoost and XGBoost reached high performance metrics individually, the ensemble approach outperformed each model by utilizing their complementary strengths. Overall, these findings underscore the need for robust machine learning methods to enhance safety standards and reliability in critical assembly processes through advanced techniques.
Tobias Gentner, Moritz Knoell, Tim Konle, Jannic Adam, Marco Klaiber, Andreas Theissler, Hermann Baumgartl
ETFA5
2025 Sky is the Limit: Exploring Solar Photovoltaic Nowcasting with Sky Images, Transfer Learning, and Data Augmentation
abstract
The increasing share of solar photovoltaic (PV) installations brings new challenges for grid stability, since the electricity generated is variable and depends on solar radiation and other meteorological factors. As previous research on nowcasting PV electricity generation has not simultaneously integrated sky images, temperature data, transfer learning, and data augmentation, we combined these approaches and evaluated various methodologies. Our machine learning models were trained on sky images and PV data derived from the SKIPP’D dataset at Stanford University and the SIRTA atmospheric observatory near Paris, enriched with temperature data from Weather Underground. The results indicate that the transfer learning model trained on augmented sky images from Stanford showed competing overall results and outperformed the current benchmark on sunny days with a root mean squared error of 0.570 on the SKIPP’D dataset. Our approach underscores the complex interactions between various predictive factors and highlights the need for sophisticated methods to enhance the nowcasting of PV power generation. Future research should explore alternative configurations and advanced architectures, such as vision transformer models, to further enhance the prediction accuracy and expand the contribution of machine learning to sustainable and environmentally friendly strategies.
Tobias Gentner, Moritz Knoell, Jannic Adam, Andreas Theissler, Marco Klaiber
KES5
2025 Simultaneous Identification and Classification of Lung Nodules in CT Images - A Hierarchical Network Approach
abstract
With ongoing advancements in Machine Learning (ML) algorithms like Convolutional Neural Networks (CNNs), their effectiveness in medical image analysis, particularly for detecting lung diseases in Computed Tomography (CT) scans, continues to improve. Although early lung nodule detection is vital, many state-of-the-art (SOTA) methods do not support simultaneous identification and classification. This separation can lead to inefficiencies and delays in clinical workflows, especially when real-time diagnosis is required. A unified approach that integrates both tasks is therefore highly desirable. Hierarchical classification models can address this by handling multi-step tasks and utilizing hierarchical label structures. This work proposes a hierarchical image classification approach for lung CT images, capable of distinguishing between nodule and non-nodule cases in the first stage and further classifying detected nodules into adenocarcinoma, small cell carcinoma, large cell carcinoma, and squamous cell carcinoma cases in the second stage. Key contributions include a method for the creation of a custom hierarchical dataset by merging the LIDC-IDRI and the Lung-PET-CT-Dx and the development of a two-level hierarchical classification algorithm with independently optimized subclassifiers. The results of this work demonstrate the effectiveness of the hierarchical model. When evaluating for simultaneous nodule identification and classification, the proposed approach achieves a weighted accuracy of 96.76%, surpassing the performance of separately executed steps. Furthermore, compared to a fat classification model, which achieves a weighted accuracy of 94.40%, the hierarchical approach outperforms it by 2.30 percentage points, highlighting the advantages of utilizing hierarchical label structures.
Dominik Hahn, Christoph Mattmann, Marco Klaiber, Sören Wagner, Marc Fernandes, Manfred Rössle
KES3
2025 A Capsule Network-Based Hybrid Model for Lung Nodule Detection
abstract
Lung cancer is a leading cause of death, where early detection is one of the only chances of survival. Manual early detection is time-consuming and error-prone, as it relies on the ability of radiologists to detect small nodules on hundreds of Computed Tomography (CT) images. To address the limitations of human analysis, Computer-Aided Detection (CAD) systems based on Convolutional Neural Networks (CNNs) can assist radiologists but require large datasets, struggle with image variations and lack interpretability. A Capsule Network (CapsNet) can offer an alternative, preserving spatial hierarchies and needing fewer training samples. This work presents a three-stage hybrid model combining You Only Look Once (YOLO) and a CapsNet for lung nodule detection and classification. The pipeline includes YOLOv11 for initial detection, a CapsNet for verification, and YOLOv11 for improved detection. The approach is evaluated on a combined LIDC-IDRI and Lung-PET-CT-Dx dataset. Experiments were done to compare the performance of a Full Image CapsNet classifier versus a Cropped Image CapsNet classifier, where only the Region of Interest (ROI) from detected nodules was used for classification. Results indicate that the Cropped Image CapsNet achieves superior performance with a classification accuracy of 95.57%, compared to 93.07% for the Full Image CapsNet. Despite improved precision, the three-stage pipeline shows slightly lower overall detection accuracy (89.98%) than standalone YOLOv11 (90.79%) due to increased False Negatives (FNs). While a CapsNet enhances classification reliability for FNs, further improvements in bounding box generation and a CapsNet architecture are needed to mitigate its higher False Positive (FP) rate. Future research should optimize the CapsNet, explore alternative detection models, and apply multi-class classification to distinguish benign from malignant nodules. The findings of this study contribute to the ongoing efforts in improving automated lung cancer diagnosis, offering a approach that leverages both CNN-based object detection and a CapsNet’s spatial feature encoding capabilities to enhance nodule detection accuracy and reduce radiologist workload.
Christoph Mattmann, Dominik Hahn, Marco Klaiber, Sören Wagner, Marc Fernandes, Manfred Rössle
KES3
2025 Dynamic Descriptive Analytics in Football: A Case Study with Retrieval-Augmented Generation for Structured Data
abstract
The rapid evolution of football (soccer) analytics has been driven by advances in structured data analysis and recently also by Large Language Models (LLMs). However, existing methods often fail to adapt dynamically to evolving queries and lack contextual richness. This paper presents a novel retrieval augmented generation (RAG) approach tailored to descriptive football analytics, which leverages spatio-temporal and opponent-related data to transform structured event and player data into actionable insights. Our approach was evaluated using a subset from the 2023/24 season of the first German division (1. Bundesliga) over multiple game weeks, achieving an average accuracy of 63.3% in generating responses, setting a first benchmark. In particular, our approach demonstrated strong performance in answering temporal and spatial queries with an accuracy of 70%, while challenges in player-specific queries highlight opportunities for further refinement. These results underscore the potential of RAG to improve decision making for analysts, sports journalists, and potentially coaches by providing dynamic and query-specific insights, paving the way for advanced applications in descriptive sports analytics and interdisciplinary approaches in digital transformation.
Ioannis Tzikas, Samuel Didovic, Felix Gerschner, Manfred Rössle, Andreas Theissler, Marco Klaiber
KES6
2024 Enhancing Website Fraud Detection: A ChatGPT-Based Approach to Phishing Detection
abstract
Phishing attacks continue to be a major cyber security problem, leading to an increasing number of studies looking at defense strategies. Therefore, we propose an LLM-based phishing detection approach that enhances work by Koide et al. by extending the prompts with URLs, adapting the Chain-of-Thought (CoT) and incorporating additional parameters. Our approach calculates a phishing score, which is used for the classification of websites as either phishing or non-phishing. Subsequently, the results of our research should enable the development of more effective LLM-based phishing detection systems and aim to improve cyber security defenses against this threat.
Michael Schesny, Nico Lutz, Thomas Jägle, Felix Gerschner, Marco Klaiber, Andreas Theissler
COMPSAC5
2024 Open-Source Text-to-Image Models: Evaluation using Metrics and Human Perception
abstract
Text-to-image models, which aim to convert text input into images, have gained popularity partly due to their flex-ibility and user-friendliness. However, there are still weaknesses in the generation of images intended to display emotions, visual text, multiple objects, relative positioning, and attribute binding. This study analyzes the weaknesses of three open-source models: Stable Diffusion v2-1, Openjourney, and Dreamlike Photoreal 2.0. The models are compared based on scores for quality, alignment, and aesthetics. The evaluation is based on (a) the metrics ClipS core, Frechet Inception Distance (FID), and Large-scale Artificial Intelligence Open Network (LAION) and (b) human perception obtained in user surveys. The evaluation revealed that all models show predominantly unsatisfactory performance, and the identified weaknesses were confirmed.
Aylin Yamac, Dilan Genc, Esra Zaman, Felix Gerschner, Marco Klaiber, Andreas Theissler
COMPSAC5
2024 Leveraging GenAI for an Intelligent Tutoring System for R: A Quantitative Evaluation of Large Language Models
abstract
The tremendous advances in Artificial Intelligence (AI) open new opportunities for education, with Intelligent Tutoring Systems (ITS) powered by Generative Artificial Intelligence (GenAI) proving to be a promising prospect. Because of this, our work explores state-of-the-art (SOTA) ITS approaches with the integration of Large Language Models (LLMs) to improve programming education. We investigate whether and how a GenAI-based ITS can effectively support students in learning R programming skills. We measured the performance of three current pairings of LLMs and user interfaces: GPT-3.5 via ChatGPT, PaLM 2 via Google Bard, and GPT-4 via Bing. Therefore, we evaluated the LLMs on four types of problem settings when learning/teaching programming. Our experimental results show that the use of generative AI, specifically LLMs for R programming, is promising, where GPT-3.5 yielded the most satisfactory results. Furthermore, the advantages and limitations of our approach are addressed and revealed. Finally, open research directions towards explainable AI (XAI) and integrated self-assessment are pointed out.
Lukas Frank, Fabian Herth, Paul Stuwe, Marco Klaiber, Felix Gerschner, Andreas Theissler
EDUCON4
2024 COVID-19 and its early Diagnosis: A Systematic Literature Review of SOTA Machine Learning Approaches
abstract
The coronavirus disease 2019 (COVID-19) has had and continues to have a major impact on public health worldwide. Therefore, early detection of COVID-19 is of great importance to control the spread of the pandemic. In the course, several Machine Learning (ML) and Deep Learning (DL) approaches using various imaging modalities such as X-ray, CT, or ultrasound images have been introduced to enable faster detection and better decision making. In this context, this paper provides an overview of the state of the art (SOTA) and the development of ML and DL in the detection of COVID-19. Based on a Systematic Literature Review (SLR), a comprehensive and systematic tabular was created with the key aspects of the identified articles, such as the image modality, the techniques used for recognition, the dataset, the number of images, and the performance metrics of the developed approach. In addition, the advantages and disadvantages of the individual approaches are discussed and the need for further research is identified, which is also intended to help fight against potential future diseases.
Sophia Kärger, Marco Klaiber, Felix Gerschner, Marc Fernandes, Manfred Rössle
KES2
2024 Monitoring Applications with Sound Data: A Systematic Literature Review on Sound Classification with Transfer Learning
abstract
Audio Classification using Machine Learning (ML) techniques has gained significant importance in various domains such as speech recognition, music Classification, and environmental sound analysis. Especially in combination with Transfer Learning (TL), this is a promising technique, which is why we conduct a Systematic Literature Review (SLR) on approaches in this domain, with a focus on sound Classification for monitoring tasks, which differ significantly from speech and music Classification. Furthermore, we provide an overview of TL techniques and applications, considering different methods due to the inherent characteristics of acoustic sound data. Based on our SLR, the advantages and disadvantages of the approaches are highlighted, and further research needs are identified.
Fabian Klärer, Jonas Werner, Marco Klaiber, Felix Gerschner, Manfred Rössle
KES3
2024 Federated Learning for Sound Data: Accurate Fan Noise Classification with a Multi-Microphone Setup
abstract
Fan sound Classification is a challenging task due to the complexity of acoustic conditions and the distinctive characteristics of different microphones. This study presents an in-depth analysis of fan sound Classification using Federated Learning (FL) across three microphone setups. We evaluate the impact of microphone variations on the performance of the five FL strategies - FedAvg, FedAdagrad, FedYogi, FedAdam, and FedMedian - and explore the potential of FL for decentralized audio Classification. Our comprehensive comparison of these strategies identifies FedYogi as the most effective, demonstrating exceptional adaptability and robustness across diverse acoustic conditions. This investigation not only sheds light on the complex dynamics between microphone variations and Classification accuracy, but also offers deep insights into the application of FL for sound-based equipment monitoring. Furthermore, our findings underscore the significant impact of microphone selection on the efficacy of FL strategies, reinforcing the need for careful consideration of hardware in the deployment of FL systems.
Kim Niklas Neuhäusler, Nico Harald Wittek, Marco Klaiber, Felix Gerschner, Marc Fernandes, Manfred Rössle
KES3
2024 Extraction of Measurement Device Information on an ESP32 Microcontroller: TinyML for Image Processing
abstract
Convolutional neural networks (CNNs) have demonstrated outstanding results in various areas of computer vision (CV). This success has led to the possibility of using CV on ever smaller computing devices, giving rise to the research area TinyML, which enables ML tasks on, e.g. resource-constrained microcontrollers. On this basis, we extend the scope of TinyML and present an image regression task where a self-generated dataset is introduced. We compare eight different approaches with different CNN architectures and normalization methods, with the best performing model achieving an MAE of 0.54 on an ESP-32. Furthermore, the ML models used are compared in terms of their performance when used on an ESP32 and on a PC. Finally, we present open questions and further research directions based on our results.
Jonas Paul, Marco Klaiber, Manfred Rössle
KES3
2023 The 10 most popular Concept Drift Algorithms: An overview and optimization potentials
abstract
In a dynamic world, data streams are continuously generated, which poses immense challenges for machine learning (ML) algorithms to adapt to changing statistical properties that are subject to a non-stationary context. The underlying scenario is defined as concept drift (CD), where changes in the relationship between response and prediction variables (real CD) or a change in input data (virtual CD) are accompanied by a significant degradation in the predictive performance of the models, causing ML models to reach unacceptable levels of system accuracy. In this paper, the state of the art for CD algorithms is analyzed and compared. For this purpose, a systematic literature review was performed. Then, the 10 most popular CD algorithms were extracted from the literature using a newly-developed metric. Subsequently, the algorithms were analyzed and compared with respect to their functionality and limitations. Based on these, the optimization potentials were systematically derived. This work presents a summarized overview of CD algorithms and provides the basis for algorithm optimization in this domain.
Marco Klaiber, Manfred Rössle, Andreas Theissler
KES1
2023 From Data to Wisdom: A Review of Applications and Data Value in the context of Small Data
abstract
Small data and big data are distinct approaches to data analysis and utilization in various applications. While big data has been the focus of many research and business efforts for more than ten years, small data is increasingly being recognized as having potential value in certain settings. We systematically review literature and conclude that small data can be indeed valuable in certain scenarios. This paper incorporates the data value perspective of small data within various application areas. For this, we apply the data-information-knowledge-wisdom (DIKW) hierarchy to categorize papers and findings, and discuss the papers from the view point of “data value”. Our review identifies various contexts where small data can be used to create value, such as data pre-processing, classification tasks, anomaly detection, forecasting and decision support. We also highlight industries that may be particularly promising areas for practitioners and researchers focused on small data. In addition, we provide an overview of methods and tools for small data analysis, including statistical techniques, visualization, and machine learning algorithms. Finally, based on our results, we suggest, that further research should focus on small data analysis.
Jonas Werner, Philipp Beisswanger, Christoph Schürger, Marco Klaiber, Andreas Theissler
KES4
2021 A Systematic Literature Review on Transfer Learning for 3D-CNNs
abstract
The dependence of convolutional neural networks on large-scale datasets for training is no secret. This is even more problematic when using 3D-CNNs since sufficient 3D datasets for training are scarce and expensive. One possible solution is transfer learning. In this comparison, the state-of-the-art techniques for 3D-CNNs are analyzed and compared. Therefore, a literature search in the databases IEEEXplore DL, ScienceDirect, SpringerLink, and ACM is conducted. The results are compared using the criteria field of application, datasets, 3D-CNN architecture, transfer learning technique, hyperparameters, and final performance. This comparison provides a basis for future work to promote understanding and usage of transfer learning for 3D-CNNs.
Marco Klaiber, Daniel Sauter, Hermann Baumgartl, Ricardo Buettner
IJCNN1