VLDB 2026 Research / reviewers in the wild / expert
Helmut Neukirchen
dblp:21/4740 · also Helmut Wolfram Neukirchen
· DBLP profile ↗
15ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-8595-3748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resource-adaptive successive doubling for hyperparameter optimization with large datasets on high-performance computing systemsabstractThe accuracy of Machine Learning (ML) models is highly dependent on the hyperparameters that have to be chosen by the user before the training. However, finding the optimal set of hyperparameters is a complex process, as many different parameter combinations need to be evaluated, and obtaining the accuracy of each combination usually requires a full training run. It is therefore of great interest to reduce the computational runtime of this process. On High-Performance Computing (HPC) systems, several configurations can be evaluated in parallel to speed up this Hyperparameter Optimization (HPO). State-of-the-art HPO methods follow a bandit-based approach and build on top of successive halving, where the final performance of a combination is estimated based on a lower than fully trained fidelity performance metric and more promising combinations are assigned more resources over time. Frequently, the number of epochs is treated as a resource, letting more promising combinations train longer. Another option is to use the number of workers as a resource and directly allocate more workers to more promising configurations via data-parallel training. This article proposes a novel Resource-Adaptive Successive Doubling Algorithm (RASDA), which combines a resource-adaptive successive doubling scheme with the plain Asynchronous Successive Halving Algorithm (ASHA). Scalability of this approach is shown on up to 1,024 Graphics Processing Units (GPUs) on modern HPC systems. It is applied to different types of Neural Networks (NNs) and trained on large datasets from the Computer Vision (CV), Computational Fluid Dynamics (CFD), and Additive Manufacturing (AM) domains, where performing more than one full training run is usually infeasible. Empirical results show that RASDA outperforms ASHA by a factor of up to 1.9 with respect to the runtime. At the same time, the solution quality of final ASHA models is maintained or even surpassed by the implicit batch size scheduling of RASDA. With RASDA, systematic HPO is applied to a terabyte-scale scientific dataset for the first time in the literature, enabling efficient optimization of complex models on massive scientific data. Marcel Aach, Rakesh Sarma, Helmut Neukirchen, Morris Riedel, Andreas Lintermann |
Future Gener. Comput. Syst. | 3 |
| 2022 | Reproducible Cross-border High Performance Computing for Scientific PortalsabstractTo reproduce eScience, several challenges need to be solved: scientific workflows need to be automated; the involved software versions need to be provided in an unambiguous way; input data needs to be easily accessible; High-Performance Computing (HPC) clusters are often involved and to achieve bit-to-bit reproducibility, it might be even necessary to execute the code on a particular cluster to avoid differences caused by different HPC platforms (and unless this is a scientist's local cluster, it needs to be accessed across (administrative) borders). Preferably, to allow even inexperienced users to (re-)produce results, all should be user-friendly. While some easy-to-use web-based scientific portals support already to access HPC resources, this typically only refers to computing and data resources that are local. By the example of two community-specific portals in the fields of biodiversity and climate research, we present a solution for accessing remote HPC (and cloud) compute and data resources from scientific portals across borders, involving rigorous container-based packaging of the software version and setup automation, thus enhancing reproducibility. Kessy Abarenkov, Anne Fouilloux, Helmut Neukirchen, Abdulrahman Azab |
e-Science | 3 |
| 2022 | Facilitating Collaboration in Machine Learning and High-Performance Computing Projects with an Interaction RoomabstractThe design, development, and deployment of scientific computing applications can be quite complex, in particular when involving Machine Learning (ML) or High-Performance Computing (HPC). They require scientific and software engineering expertise and in addition HPC or ML knowledge. Often, such applications are however developed by scientists who are experts in their domain, but need support for the software engineering, ML, and HPC aspects. The cooperation and communication between experts from these quite different disciplines can be difficult though. We therefore propose to employ the Interaction Room (IR), a method that facilitates interdisciplinary collaboration in complex software projects. An IR uses annotated drawings to exchange information and stimulate discussion between project stakeholders, in order to improve common understanding and identify uncertainties, risks, and other aspects that are critical to a project's success early on. We suggest different drawing canvases and annotations that focus on different viewpoints and issues of the project. These canvases are specific to the project type, such as ML applications or HPC simulations. Matthias Book, Morris Riedel, Helmut Neukirchen, Ernir Erlingsson |
e-Science | 3 |
| 2022 | Accelerating Hyperparameter Tuning of a Deep Learning Model for Remote Sensing Image ClassificationabstractDeep Learning models have proven necessary in dealing with the challenges posed by the continuous growth of data volume acquired from satellites and the increasing complexity of new Remote Sensing applications. To obtain the best performance from such models, it is necessary to fine-tune their hyperparameters. Since the models might have massive amounts of parameters that need to be tuned, this process requires many computational resources. In this work, a method to accelerate hyperparameter optimization on a High-Performance Computing system is proposed. The data batch size is increased during the training, leading to a more efficient execution on Graphics Processing Units. The experimental results confirm that this method reduces the runtime of the hyperparameter optimization step by a factor of 3 while achieving the same validation accuracy as a standard training procedure with a fixed batch size. Marcel Aach, Rocco Sedona, Andreas Lintermann, Gabriele Cavallaro, Helmut Neukirchen, Morris Riedel |
IGARSS | 5 |
| 2019 | Scalable Workflows for Remote Sensing Data Processing with the Deep-Est Modular Supercomputing ArchitectureabstractThe implementation of efficient remote sensing workflows is essential to improve the access to and analysis of the vast amount of sensed data and to provide decision-makers with clear, timely, and useful information. The Dynamical Exascale Entry Platform (DEEP) is an European pre-exascale platform that incorporates heterogeneous High-Performance Computing (HPC) systems, i.e., hardware modules which include specialised accelerators. This paper demonstrates the potential of such diverse modules for the deployment of remote sensing data workflows that include diverse processing tasks. Particular focus is put on pipelines which can use the Network Attached Memory (NAM), which is a novel supercomputer module that allows near processing and/or fast shared storage of big remote sensing datasets. Ernir Erlingsson, Gabriele Cavallaro, Helmut Neukirchen, Morris Riedel |
IGARSS | 3 |
| 2018 | Scaling Support Vector Machines Towards Exascale Computing for Classification of Large-Scale High-Resolution Remote Sensing ImagesabstractProgress in sensor technology leads to an ever-increasing amount of remote sensing data which needs to be classified in order to extract information. This big amount of data requires parallel processing by running parallel implementations of classification algorithms, such as Support Vector Machines (SVMs), on High-Performance Computing (HPC) clusters. Tomorrow's supercomputers will be able to provide exascale computing performance by using specialised hardware accelerators. However, existing software processing chains need to be adapted to make use of the best fitting accelerators. To address this problem, a mapping of an SVM remote sensing classification chain to the Dynamical Exascale Entry Platform (DEEP), a European pre-exascale platform, is presented. It will allow to scale SVM-based classifications on tomorrow's hardware towards exascale performance. Ernir Erlingsson, Gabriele Cavallaro, Morris Riedel, Helmut Neukirchen |
IGARSS | 4 |
| 2018 | Automated Analysis of Remotely Sensed Images Using the Unicore Workflow Management SystemabstractThe progress of remote sensing technologies leads to increased supply of high-resolution image data. However, solutions for processing large volumes of data are lagging behind: desktop computers cannot cope anymore with the requirements of macro-scale remote sensing applications; therefore, parallel methods running in High-Performance Computing (HPC) environments are essential. Managing an HPC processing pipeline is non-trivial for a scientist, especially when the computing environment is heterogeneous and the set of tasks has complex dependencies. This paper proposes an end-to-end scientific workflow approach based on the UNICORE workflow management system for automating the full chain of Support Vector Machine (SVM)-based classification of remotely sensed images. The high-level nature of UNICORE workflows allows to deal with heterogeneity of HPC computing environments and offers powerful workflow operations such as needed for parameter sweeps. As a result, the remote sensing workflow of SVM-based classification becomes re-usable across different computing environments, thus increasing usability and reducing efforts for a scientist. M. Shahbaz Memon, Gabriele Cavallaro, Björn Hagemeier, Morris Riedel, Helmut Neukirchen |
IGARSS | 5 |
| 2018 | Towards Federated Service Discovery and Identity Management in Collaborative Data and Compute Cloud Infrastructures
Ahmed Shiraz Memon, Jens Jensen, Willem Elbers, Helmut Neukirchen, Matthias Book, Morris Riedel |
J. Grid Comput. | 4 |
| 2017 | Facilitating efficient data analysis of remotely sensed images using standards-based parameter sweep modelsabstractClassification of remote sensing images often use Support Vector Machines (SVMs) that require an n-fold cross-validation phase in order to do model selection. This phase is characterized by sweeping through a wide set of parameter combinations of SVM kernel and cost parameters. As a consequence this process is computationally expensive but represents a principled way of tuning a model for better accuracy and to prevent overfitting together with regularization that is in SVMs inherently solved in the optimization. Since the cross-validation technique is done in a principled way also known as `gridsearch', we aim at supporting remote sensing scientists in two ways. Firstly by reducing the time-to-solution of the cross-validation by applying state-of-the-art parallel processing methods because the sweep of parameters and cross-validation runs itself can be nicely parallelized. Secondly by reducing manual labour by automating the parallel submission processes since manually performing cross-validation is very time consuming, unintuitive, and error-prone especially in large-scale cluster or supercomputing environments (e.g., batch job scripts, node/core/task parameters, etc.). M. Shahbaz Memon, Gabriele Cavallaro, Morris Riedel, Helmut Neukirchen |
IGARSS | 4 |
| 2011 | An empirical study of software architectures' effect on product quality
Klaus Marius Hansen, Kristjan Jonasson, Helmut Neukirchen |
J. Syst. Softw. | 3 |
| 2009 | A Flexible Framework for Quality Assurance of Software Artefacts with Applications to Java, UML, and TTCN-3 Test SpecificationsabstractManual reviews and inspections of software artefacts are time consuming and thus, automated analysis tools have been developed to support the quality assurance of software artefacts. Usually, software analysis tools are implemented for analysing only one specific language as target and for performing only one class of analyses. Furthermore, most software analysis tools support only common programming languages, but not those domain-specific languages that are used in a test process. As a solution, a framework for software analysis is presented that is based on a flexible, yet high-level facade layer that mediates between analysis rules and the underlying target software artefact; the analysis rules are specified using high-level XQuery expressions. Hence, further rules can be quickly added and new types of software artefacts can be analysed without needing to adapt the existing analysis rules. The applicability of this approach is demonstrated by examples from using this framework to calculate metrics and detect bad smells in Java source code, in UML models, and in test specifications written using the Testing and Test Control Notations (TTCN-3). Jens Nodler, Helmut Neukirchen, Jens Grabowski |
ICST | 2 |
| 2008 | Testing Grid Application Workflows Using TTCN-3abstractThe collective and coordinated usage of distributed resources for problem solution within dynamic virtual organizations can be realized with the grid computing technology. For distributing and solving a task, a grid application involves a complex workflow of dividing a task into smaller sub-tasks, scheduling and submitting jobs for solving those sub-tasks, and eventually collecting and combining the results of the sub-tasks into a final result. The quality assurance of grid applications is a challenge due to the highly distributed nature of the grid environment in which the grid application is deployed. This paper investigates the applicability of the testing and test control notation (TTCN-3) for testing the workflows of distributed grid applications. To this aim, a case study has been created that consists of a distributed grid application which includes a typical grid application workflow; as the main contribution, this case study contains a corresponding distributed TTCN-3 test suite that tests the correct execution of the grid application workflow. To demonstrate the adaptation of the abstract TTCN-3 test suite to a specific grid environment, corresponding reusable test adapters have been implemented for the grid middleware Globus Toolkit 4 (GT4). The realized test system demonstrates that TTCN-3 is applicable for testing the workflow of distributed grid applications. Thomas Rings, Helmut Neukirchen, Jens Grabowski |
ICST | 2 |
| 2008 | An approach to quality engineering of TTCN-3 test specificationsabstractExperience with the development and maintenance of large test suites specified using the Testing and Test Control Notation (TTCN-3) has shown that it is difficult to construct tests that are concise with respect to quality aspects such as maintainability or usability. The ISO/IEC standard 9126 defines a general software quality model that substantiates the term “quality” with characteristics and subcharacteristics. The domain of test specifications, however, requires an adaption of this general model. To apply it to specific languages such as TTCN-3, it needs to be instantiated. In this paper, we present an instantiation of this model as well as an approach to assess and improve test specifications. The assessment is based on metrics and the identification of code smells. The quality improvement is based on refactoring. Example measurements using our TTCN-3 tool TRex demonstrate how this procedure is applied in practise. Helmut Neukirchen, Benjamin Zeiss, Jens Grabowski |
Int. J. Softw. Tools Technol. Transf. | 1 |
| 2008 | Quality assurance for TTCN-3 test specificationsabstractAbstract Comprehensive testing of modern communication systems often requires large and complex test suites, which have to be maintained throughout the system life cycle. Industrial experience, with those written using the standardized Testing and Test Control Notation (TTCN‐3), has shown that this maintenance is a non‐trivial task and its burden can be reduced by means of appropriate concepts and tool support. To this aim, Motorola has collaborated with the University of Göttingen to develop TRex, an open‐source TTCN‐3 development environment, which notably provides suitable metrics and refactorings to enable the assessment and automatic restructuring of test suites. This article presents concepts like metrics and refactoring for the quality assurance of TTCN‐3 test suites and their implementation provided by the TRex tool. These means make it far easier to construct and maintain TTCN‐3 tests that are concise and optimally balanced with respect to maintainability quality characteristics. Copyright © 2008 John Wiley & Sons, Ltd. Helmut Neukirchen, Benjamin Zeiss, Jens Grabowski, Paul Baker, Dominic Evans |
Softw. Test. Verification Reliab. | 1 |
| 2000 | Managing Services in Distributed Systems by Integrating Trading and Load BalancingabstractWith a changing structure of networks and application systems due to the requirements of decentralised enterprises and open service markets, distributed systems with rapidly increasing complexity are evolving. New concepts for an efficient management of such systems have to be developed. Focusing on the service level, examples for existing concepts are trading to find services in a distributed environment, and load balancing to avoid performance bottlenecks in service provision. This paper discusses the integration of a load balancer into a trader to adapt the allocation of client requests to suitable servers due to the current system usage, and thus to improve the quality of the services in terms of performance. The approach used is independent of the servers' characteristics, because for the servers involved, no provision of additional service properties to cover load aspects is necessary. Furthermore, it is flexible to enhance, because the concept of load used can be varied without modification of trader or load balancer. Dirk Thißen 0001, Helmut Neukirchen |
ISCC | 2 |