EDBT 2026 Demo / reviewers in the wild / expert
Siegfried Benkner
dblp:s/SiegfriedBenkner
· DBLP profile ↗
53ranked-venue papers
17as first author
9since 2021 · last 2025
0000-0002-6520-2047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 11 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorArtificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Reactive Autoscaling for Task-Based Pipelines on KubernetesabstractWe present Python-to-Kubernetes (PTK), a hybrid autoscaling framework for pipeline-oriented, task-based Python applications on Kubernetes. PTK coordinates queue-length-driven horizontal scaling for CPU, memory, and GPU, together with reactive in-place vertical scaling of CPU and memory. The framework introduces source-code annotations, enabling users to define task-specific scaling constraints and automatically generate Kubernetes manifests. A periodic controller uses utilization and queue metrics to coordinate horizontal and vertical scaling, improving resource efficiency while maintaining pipeline performance. In a streaming machine learning (ML) inference pipeline, PTK sustains the target throughput while reducing hourly cost by 40.6%, CPU by 32.1 %, and memory by 22.4%, and lowering the GPU count from 4 to 3, compared with an uncoordinated baseline that combines the Horizontal Pod Autoscaler (HPA) and the Vertical Pod Autoscaler (VPA). It also cuts peak cost by 23.6% compared with a queue-driven HPA baseline. Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner |
CloudCom | 3 |
| 2025 | Deadline-Aware Resource Allocation and Scheduling of Serverless Workloads on Heterogeneous ClustersabstractServerless computing has become widely adopted as a cloud deployment model due to its ease of use and finegrained pay-as-you-go pricing. By hiding infrastructure complexity, it simplifies access to cloud resources and lets developers focus on application code. However, most serverless platforms operate on a best-effort basis and provide minimal control over performance tuning. Combined with limited visibility into underlying hardware, this makes it difficult to reliably meet Service Level Objectives (SLOs). To address this, we introduce DHRT, a deadline- and heterogeneity-aware scheduling and resource allocation framework for performance-critical serverless workloads. DHRT applies heuristic-driven online optimisation to iteratively refine resource estimates by leveraging real-time metrics and historical data from live executions. To fulfil SLOs, it accounts for both workload characteristics and node heterogeneity. We evaluate DHRT on synthetic workloads by comparing it against baseline scheduling and resource allocation policies commonly used in FaaS platforms. Results show that DHRT accurately estimates resource demands within a few live executions, eliminating the need for manual resource tuning. By exploiting node heterogeneity and dynamically scaling vCPU allocations as workloads near their deadlines, DHRT improves resource efficiency and significantly reduces deadline violations. Matthias Fritz, Siegfried Benkner, Enes Bajrovic |
CLUSTER | 2 |
| 2025 | Provisioning of Kubernetes Clusters for Task-Based Python ApplicationsabstractWe present Python-to-Kubernetes (PTK), a framework that automates the provisioning and deployment of taskbased Python applications on Kubernetes. PTK introduces compact source-code annotations for tasks, resource needs, grouping, and data-size hints. From these annotations, it provisions an application-specific cluster, builds container images, generates manifests, and selects the data-transfer mechanism based on placement. A scoring-based mapping co-locates bandwidth-heavy neighbors to reduce cross-node traffic and right-sizes nodes after placement. On a six-task machine learning (ML) ResNet50 imageclassification pipeline (ImageNet$\mathbf{5 \%} \boldsymbol{/} \mathbf{1 0 \%}$subsets), PTK achieves up to$6.43 \times$faster runtime and$7.26 \times$lower cost per run than the Kubernetes Default Scheduler, with higher CPU/memory utilization and fewer/smaller nodes. These results indicate that lightweight annotations plus application-aware provisioning can substantially improve price-performance for Kubernetes-based ML pipelines. Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner |
ICPADS | 3 |
| 2025 | Accelerating Graph Neural Networks Using a Novel Computation-Friendly Matrix Compression FormatabstractThis paper proposes the Compressed Binary Matrix (CBM) format, a novel, computation-friendly compression scheme for binary matrices. CBM not only reduces the memory footprint of the matrix but also enables faster matrix multiplication between binary and dense, real-valued matrices. The CBM format can be applied to accelerate various graph-related tasks, where the (binary) adjacency matrix of the graph is repeatedly multiplied by another matrix, such as during inference and training of various types of Graph Neural Networks (GNNs). The format is evaluated on a shared-memory architecture in both serial and parallel settings. Experimental results show that CBM can reduce the memory footprint of real-world graphs up to$11 \times$, and that the parallel matrix multiplication using CBM is more than$5 \times$faster than state-of-the-art sparse-dense matrix multiplication kernels. Furthermore, when applied to the inference stage of Graph Convolutional Networks (GCNs), the CBM format achieves speedups close to$2.5 \times$compared to inference using other parallel matrix multiplication kernels. João Nuno Ferreira Alves, Samir Moustafa, Siegfried Benkner, Alexandre P. Francisco, Wilfried N. Gansterer, Luís M. S. Russo |
IPDPS | 3 |
| 2024 | Python to Kubernetes: A Programming and Resource Management Framework for Compute-and Data-intensive ApplicationsabstractIn this paper, we introduce the Python to Kubernetes (PTK) framework, a high-level Python-based programming framework for deploying Python applications on top of Kubernetes clusters. PTK supports a task-based programming approach with extensions for specifying resource requirements and performance constraints. A major goal of PTK is to provide users with high-level control for deploying compute- and data-intensive applications on different types and configurations of heterogeneous clusters, while ensuring performance and/or cost constraints. Andrey Nagiyev, Enes Bajrovic, Siegfried Benkner |
ICPADS | 3 |
| 2023 | Experiences in Architectural Design and Deployment of eHealth and Environmental Applications for Cloud-Edge Continuum
Atakan Aral, Antonio Esposito 0001, Andrey Nagiyev, Siegfried Benkner, Beniamino Di Martino, Mario A. Bochicchio |
AINA (3) | 4 |
| 2023 | A Novel Triangular Space-Filling Curve for Cache-Oblivious In-Place Transposition of Square MatricesabstractThis paper proposes a novel cache-oblivious blocking scheme based on a new triangular space-filling curve which preserves data locality. The proposed blocking-scheme reduces the movement of data within the host memory hierarchy for triangular matrix traversals, which inherently exhibit poor data locality, such as the in-place transposition of square matrices. We show that our cache-oblivious blocking-scheme can be generated iteratively in linear time and constant memory with regard to the number of entries present in the lower, or upper, triangle of the input matrix. In contrast to classical recursive cache-oblivious solutions, the iterative nature of our blocking-scheme does not inhibit other essential optimizations such as software prefetching. In order to assess the viability of our blocking-scheme as a cache-oblivious strategy, we applied it to the in-place transposition of square matrices. Extensive experiments show that our cache-oblivious transposition algorithm generally outperforms the cache-aware state-of-the-art algorithm in terms of throughput and energy efficiency in sequential as well as parallel environments. João Nuno Ferreira Alves, Luís M. S. Russo, Alexandre P. Francisco, Siegfried Benkner |
IPDPS | 4 |
| 2022 | The OCR-Vx experience: lessons learned from designing and implementing a task-based runtime systemabstractTask-based runtime systems are an important branch of parallel programming research, since tasks decouple computation from the compute units, giving the runtime systems greater flexibility than a thread-based solution. This makes it easier to deal with the ever-increasing complexity of parallel architectures by providing a separation of concerns-the specification of parallelism is separated from the implementation of the parallel computations on a specific architecture. The Open Community Runtime is one such system, aimed at large-scale parallel systems. Unlike many other task-based runtime systems, the creators not only provided an implementation but there is also a comprehensive specification document. This has allowed us to create an independent implementation, called OCR-Vx. In this article, we present our experience of developing the runtime system, put our work in the context of the specification and the other implementations, and describe key lessons that we have learned during our work. We discuss the design and implementation issues of task-based runtime systems and applications including task synchronization and scheduling, data management, memory consistency, the relation between shared-memory and distributed-memory runtime systems, NUMA architectures, and heterogeneous systems. The article is aimed at audiences not familiar with OCR, since we believe these lessons could be valuable for developers working on other task-based runtime systems or designing new ones. Jirí Dokulil, Siegfried Benkner |
J. Supercomput. | 2 |
| 2021 | Matching Program Implementations and Heterogeneous Computing Systems
Martin Sandrieser, Siegfried Benkner |
PDCAT | 2 |
| 2020 | Automatic Placement of Tasks to NUMA Nodes in Iterative ApplicationsabstractManycore architectures with non-uniform memory access (NUMA) are commonly used for high-performance computing. On these systems, the placement of data and computation to the NUMA nodes has very significant impact on performance, especially with memory-bound applications. This placement can usually be defined by the programmer, but it is generally desirable to automate the placement to simplify the programmer's job, improve portability, and make the code more future-proof. Task-based runtime systems already assume a fair degree of responsibility for task placement, so it is only natural to involve them in mapping work and data to NUMA nodes. In this work, we propose a solution where the runtime system first performs a profiling run of the application and measures various performance characteristics. Then, the data collected by the profiling run is used by a stand-alone analyzer to create a plan for placing tasks to NUMA nodes so that the tasks are close to the data that they most rely on. This plan is then used by the runtime system to execute the application more efficiently. We focus on iterative applications, where the same patterns of tasks are being repeated. We identify these patterns and use them to create a plan that works for any number of iterations, not just the one that used when the application was observed. In our experiments, which were performed on modern manycore systems (Intel Skylake and AMD Zen) with 4 and 8 NUMA nodes, the proposed automated placement can either match or come close (<; 10$%) to a hand-tuned placement. Jirí Dokulil, Siegfried Benkner |
PDP | 2 |
| 2020 | A benchmark set of highly-efficient CUDA and OpenCL kernels and its dynamic autotuning with Kernel Tuning Toolkit
Filip Petrovic, David Strelák, Jana Hozzová, Jaroslav Olha, Richard Trembecký, Siegfried Benkner, Jiri Filipovic |
Future Gener. Comput. Syst. | 6 |
| 2020 | Programming languages for data-Intensive HPC applications: A systematic mapping study
Vasco Amaral 0001, Beatriz Norberto, Miguel Goulão, Marco Aldinucci, Siegfried Benkner, Andrea Bracciali, Paulo Carreira 0001, Edgars Celms, Luís Correia 0001, Clemens Grelck, Helen D. Karatza, Christoph W. Kessler, Peter Kilpatrick, Hugo F. M. C. Martiniano, Ilias Mavridis, Sabri Pllana, Ana Respício, José Simão, Luís Veiga, Ari Visa |
Parallel Comput. | 5 |
| 2018 | Pipeline Patterns on Top of Task-Based Runtimes
Enes Bajrovic, Siegfried Benkner, Jirí Dokulil |
PDCAT | 2 |
| 2018 | Adaptive Scheduling of Collocated Applications Using a Task-Based Runtime SystemabstractTask-based runtime systems are considered as one of the options for dealing with the challenges of upcoming parallel architectures. The greater flexibility of these runtime systems can also be used to dynamically adjust the resources allocated to the applications, adapting to the current load of the system and the progress of the applications. In our work, we have extended our implementation of the Open Community Runtime to support dynamic adjustment of execution threads. The runtimes communicate with an agent process, which collects performance data, computes thread allocation, and instructs the runtimes to make the required adjustments. We have tested our solution under different scenarios, focusing on producer-consumer applications, where the dynamic resource management was used to keep the applications in sync, improving the overall performance in some cases. Jirí Dokulil, Siegfried Benkner |
SBAC-PAD | 2 |
| 2018 | A multi-aspect online tuning framework for HPC applications
Michael Gerndt, Siegfried Benkner, Eduardo César, Carmen B. Navarrete, Enes Bajrovic, Jirí Dokulil, Carla Guillén, Robert Mijakovic, Anna Sikora |
Softw. Qual. J. | 2 |
| 2017 | The Open Community Runtime on the Intel Knights Landing Architecture
Jirí Dokulil, Siegfried Benkner, Jakub Yaghob |
ICA3PP | 2 |
| 2016 | Implementing the Open Community Runtime for Shared-Memory and Distributed-Memory SystemsabstractThe extreme scale, complexity and performance variability of future high performance computing systems pose many new challenges to parallel programming models and runtime systems. The Open Community Runtime (OCR) is a recent effort for a task-based runtime system for extreme scale parallel systems. We have implemented the OCR specification in a shared-memory environment on top of TBB, providing an alternative to the implementation created by the OCR consortium. We have created an experimental extension that supports parallel accelerators programmed with OpenCL. We also have an implementation that targets distributed-memory systems. Despite being in an early stage of development, our implementations can achieve reasonable performance with some applications. We describe the main aspects of our OCR implementations and report on early experimental results on shared-memory and distributed-memory systems. Jirí Dokulil, Martin Sandrieser, Siegfried Benkner |
PDP | 3 |
| 2015 | OpenCL Kernel Fusion for GPU, Xeon Phi and CPUabstractKernel fusion is an optimization method, in which the code from several kernels is composed to create a new, fused kernel. It can push the performance of kernels beyond limits given for their isolated, unfused form. In this paper, we introduce a classification of different types of kernel fusion for both data dependent and data independent kernels. We study kernel fusion on three types of OpenCL devices: GPU, Xeon Phi and CPU. Those hardware platforms have quite different properties, thus, kernel fusion often affects performance in quite different ways. We analyze the impact of kernel fusion on those hardware platforms and show how it can be used to improve performance. Based on our study we also introduce a basic transformation method for generating fused kernels, which has good potential to be automatized. Jiri Filipovic, Siegfried Benkner |
SBAC-PAD | 2 |
| 2014 | Automatic Tuning of a Parallel Pattern Library for Heterogeneous Systems with Intel Xeon PhiabstractPattern libraries are important tools for high productivity application development. Their struggle for best performance is complicated by the fact that they are used to execute user-provided code, which is not known during their creation. This makes pattern libraries good candidate for automatic software tuning. In this paper, we deal with automatic online parameter tuning of the HyPHI hybrid pattern library for heterogeneous systems equipped with the Intel Xeon Phi coprocessors. We propose a framework that can be used to combine a pattern library with an existing tuning library in a practical and efficient way. Our experiments show that tuning can noticeably improve the performance of the library and it introduces very little overhead. Jirí Dokulil, Siegfried Benkner |
ISPA | 2 |
| 2014 | Towards High-Level Parallel Patterns in OpenCLabstractParallel pattern libraries (e.g., Intel TBB) are popular and useful tools for developing applications in SMP environments at a higher level of abstraction. Such libraries execute user-provided code efficiently on shared memory parallel architectures in accordance with well-defined execution patterns like parallel for-loops or pipelines. For heterogeneous architectures comprised of CPUs and accelerators, OpenCL has gained a lot of momentum. Since accelerated architectures do not provide a shared memory, it is not possible to directly use the approach taken in pattern libraries for SMP systems for OpenCL as well. In this paper, we are exploring issues and opportunities encountered by attempts to provide such patterns in the context of OpenCL. Based on a set of experiments with a scientific application on diverse OpenCL devices, we point out major pitfalls and insights, and outline directions for further efforts in developing pattern libraries for OpenCL. Jirí Dokulil, Siegfried Benkner |
PDCAT | 2 |
| 2013 | HyPHI - Task Based Hybrid Execution C++ Library for the Intel Xeon Phi CoprocessorabstractThe Intel Threading Building Blocks (TBB) C++ library introduced task parallelism to a wide audience of application developers. The library is easy to use and powerful, but it is limited to shared-memory machines. In this paper we present HyPHI, a novel library for the Intel Xeon Phi coprocessor for building applications which execute using a hybrid parallel model that exploits parallelism across host CPUs and Xeon Phi coprocessors simultaneously. Our library currently provides hybrid for-each and map-reduce. It hides the details of parallelization, work distribution and computation offloading from users while using internally TBB as its foundation. Despite the higher level of abstraction provided by our library we show that for certain types of applications we outperform codes that rely on the built-in offload support currently provided by the Intel compiler. We have performed a set of experiments with the library and created guidelines that help the developers decide in which situations they should use the HyPHI library. Jirí Dokulil, Enes Bajrovic, Siegfried Benkner, Martin Sandrieser, Beverly Bachmayer |
ICPP | 3 |
| 2013 | A Secure and Flexible Data Infrastructure for the VPH-Share CommunityabstractThe European VPH-Share project develops a comprehensive service framework with the objective of sharing clinical data, information, models and workflows focusing on the analysis of the human physiopathology within the Virtual Physiological Human (VPH) community. The project envisions an extensive and dynamic data infrastructure built on top of a secure hybrid Cloud environment. This paper presents the data service provisioning framework that builds up the data infrastructure, focusing on the deployment of data integration services in the hybrid Cloud, the associated mechanism for securing access to patient-specific datasets, and performance results for different deployment scenarios relevant within the scope of the project. Siegfried Benkner, Yuriy Kaniovskyi, Chris Borckholder, Marian Bubak, Piotr Nowakowski, Dario Ruiz Lopez, Steven Wood |
PDCAT | 1 |
| 2013 | Improving Blocking Operation Support in Intel TBBabstractThe Intel Threading Building Blocks (TBB) template library has become a popular tool for programming many-core systems. However, it is not suitable in situations where a large number of potentially blocking calls has to be made to handle long-running operations like disk access or remote data access. We have designed and implemented an add-on for the TBB that allows developers to better integrate long-running operations into their applications. We have extended TBB's task dependencies to also include blocking operations and implemented a run-time that efficiently manages these dependencies. Jirí Dokulil, Siegfried Benkner, Martin Sandrieser |
PDCAT | 2 |
| 2012 | Programmability and performance portability aspects of heterogeneous multi-/manycore systemsabstractWe discuss three complementary approaches that can provide both portability and an increased level of abstraction for the programming of heterogeneous multicore systems. Together, these approaches also support performance portability, as currently investigated in the EU FP7 project PEPPHER. In particular, we consider (1) a library-based approach, here represented by the integration of the SkePU C++ skeleton programming library with the StarPU runtime system for dynamic scheduling and dynamic selection of suitable execution units for parallel tasks; (2) a language-based approach, here represented by the Offload-C++ high-level language extensions and Offload compiler to generate platform-specific code; and (3) a component-based approach, specifically the PEPPHER component system for annotating user-level application components with performance metadata, thereby preparing them for performance-aware composition. We discuss the strengths and weaknesses of these approaches and show how they could complement each other in an integrational programming framework for heterogeneous multicore systems. Christoph W. Kessler, Usman Dastgeer, Samuel Thibault, Raymond Namyst, Andrew Richards, Uwe Dolinsky, Siegfried Benkner, Jesper Larsson Träff, Sabri Pllana |
DATE | 7 |
| 2012 | High-Level Support for Pipeline Parallelism on Many-Core Architectures
Siegfried Benkner, Enes Bajrovic, Erich Marth, Martin Sandrieser, Raymond Namyst, Samuel Thibault |
Euro-Par | 1 |
| 2012 | Using explicit platform descriptions to support programming of heterogeneous many-core systems
Martin Sandrieser, Siegfried Benkner, Sabri Pllana |
Parallel Comput. | 2 |
| 2010 | Multicore and Manycore Programming
Beniamino Di Martino, Fabrizio Petrini, Siegfried Benkner, Kirk W. Cameron, Dieter Kranzlmüller, Jakub Kurzak, Davide Pasetto, Jesper Larsson Träff |
Euro-Par (2) | 3 |
| 2010 | Supporting Molecular Modeling Workflows within a Grid Services Cloud
Martin Koehler, Matthias Ruckenbauer, Ivan Janciak, Siegfried Benkner, Hans Lischka, Wilfried N. Gansterer |
ICCSA (4) | 4 |
| 2010 | @neurIST: Infrastructure for Advanced Disease Management Through Integration of Heterogeneous Data, Computing, and Complex Processing ServicesabstractThe increasing volume of data describing human disease processes and the growing complexity of understanding, managing, and sharing such data presents a huge challenge for clinicians and medical researchers. This paper presents the @neurIST system, which provides an infrastructure for biomedical research while aiding clinical care, by bringing together heterogeneous data and complex processing and computing services. Although @neurIST targets the investigation and treatment of cerebral aneurysms, the system's architecture is generic enough that it could be adapted to the treatment of other diseases. Innovations in @neurIST include confining the patient data pertaining to aneurysms inside a single environment that offers clinicians the tools to analyze and interpret patient data and make use of knowledge-based guidance in planning their treatment. Medical researchers gain access to a critical mass of aneurysm related data due to the system's ability to federate distributed information sources. A semantically mediated grid infrastructure ensures that both clinicians and researchers are able to seamlessly access and work on data that is distributed across multiple sites in a secure way in addition to providing computing resources on demand for performing computationally intensive simulations for treatment planning and research. Siegfried Benkner, Antonio Arbona, Guntram Berti, Alessandro Chiarini, Robert Dunlop, Gerhard Engelbrecht, Alejandro F. Frangi, Christoph M. Friedrich, S. Hanser, Peer Hasselmeyer, Rod D. Hose, Jimison Iavindrasana, Martin Koehler, Luigi Lo Iacono, Guy Lonsdale, Rodolphe Meyer, Bob Moore, Hariharan Rajasekaran, Paul E. Summers, Alexander Wöhrer, Steven Wood |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2008 | @neurIST - Towards a System Architecture for Advanced Disease Management through Integration of Heterogeneous Data, Computing, and Complex Processing ServicesabstractThis paper presents the system architecture of the @neurIST project, which aims at supporting the research and treatment of cerebral aneurysms by bringing together heterogeneous data, computing and complex processing services. The architecture is generic enough to adapt it to the treatment of other diseases beyond cerebral aneurysms. The paper describes the generic requirements of the system and presents the architecture, applications and middleware technologies used to realise the system and highlights the innovations in @neurIST. Hariharan Rajasekaran, Luigi Lo Iacono, Peer Hasselmeyer, Jochen Fingberg, Paul E. Summers, Siegfried Benkner, Gerhard Engelbrecht, Antonio Arbona, Alessandro Chiarini, Christoph M. Friedrich, Martin Hofmann-Apitius, Kai Kumpf, Bob Moore, Philippe Bijlenga, Jimison Iavindrasana, Henning Müller, Rod D. Hose, Robert Dunlop, Alejandro F. Frangi |
CBMS | 6 |
| 2008 | Hybrid Performance Modeling and Prediction of Large-Scale Computing SystemsabstractPerformance is a key feature of large-scale computing systems. However, the achieved performance when a certain program is executed is significantly lower than the maximal theoretical performance of the large-scale computing system. The model-based performance evaluation may be used to support the performance-oriented program development for large-scale computing systems. In this paper we present a hybrid approach for performance modeling and prediction of parallel and distributed computing systems, which combines mathematical modeling and discrete-event simulation. We use mathematical modeling to develop parameterized performance models for components of the system. Thereafter, we use discrete-event simulation to describe the structure of system and the interaction among its components. As a result, we obtain a high-level performance model, which combines the evaluation speed of mathematical models with the structure awareness and fidelity of the simulation model. We evaluate empirically our approach with a real-world material science program that comprises more than 15,000 lines of code. Sabri Pllana, Siegfried Benkner, Fatos Xhafa, Leonard Barolli |
CISIS | 2 |
| 2008 | Grid Services for Parallel Molecular Dynamics with NAMD and CHARMM
Siegfried Benkner, Mária Lucká, Othmar Steinhauser |
ICCSA (1) | 1 |
| 2008 | Specification, planning, and execution of QoS-aware Grid workflows within the Amadeus environmentabstractAbstract Commonly, at a high level of abstraction Grid applications are specified based on the workflow paradigm. However, majority of Grid workflow systems either do not support Quality of Service (QoS), or provide only partial QoS support for certain phases of the workflow lifecycle. In this paper we present Amadeus, which is a holistic service‐oriented environment for QoS‐aware Grid workflows. Amadeus considers user requirements, in terms of QoS constraints, during workflow specification, planning, and execution. Within the Amadeus environment workflows and the associated QoS constraints are specified at a high level using an intuitive graphical notation. A distinguishing feature of our system is the support of a comprehensive set of QoS requirements, which considers in addition to performance and economical aspects also legal and security aspects. A set of QoS‐aware service‐oriented components is provided for workflow planning to support automatic constraint‐based service negotiation and workflow optimization. For improving the efficiency of workflow planning we introduce a QoS‐aware workflow reduction technique. Furthermore, we present our static and dynamic planning strategies for workflow execution in accordance with user‐specified requirements. For each phase of the workflow lifecycle we experimentally evaluate the corresponding Amadeus components. Copyright © 2007 John Wiley & Sons, Ltd. Ivona Brandic, Sabri Pllana, Siegfried Benkner |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Performance Modeling and Prediction of Parallel and Distributed Computing Systems: A Survey of the State of the ArtabstractPerformance is one of the key features of parallel and distributed computing systems. Therefore, in the past a significant research effort was invested in the development of approaches for performance modeling and prediction of parallel and distributed computing systems. In this paper we identify the trends, contributions, and drawbacks of the state of the art approaches. We describe a wide range of the performance modeling approaches that spans from the high-level mathematical modeling to the detailed instruction-level simulation. For each approach we describe how the program and machine are modeled and estimate the model development and evaluation effort, the efficiency, and the accuracy. Furthermore, we present an overall evaluation of the presented approaches Sabri Pllana, Ivona Brandic, Siegfried Benkner |
CISIS | 3 |
| 2007 | Nonadiabatic Ab Initio Surface-Hopping Dynamics Calculation in a Grid Environment - First Experiences
Matthias Ruckenbauer, Ivona Brandic, Siegfried Benkner, Wilfried N. Gansterer, Osvaldo Gervasi, Mario Barbatti, Hans Lischka |
ICCSA (1) | 3 |
| 2007 | Component-oriented application construction for a Web service-based GridabstractAbstract We present the architecture and prototype implementation of a component‐oriented programming environment for a Web service based computational Grid. As middleware, we utilize the Vienna Grid Environment (VGE), a framework that enables the provision of compute‐intensive parallel applications as configurable, QoS‐aware Grid services. Our component model follows the Common Component Architecture (CCA) and models application Web services as distributed components. We describe a component framework that integrates VGE services with a component model allowing to express and dynamically manage application and performance meta‐data as well as dependencies on the infrastructure or other components. Furthermore, we show how the client programming interface is used to compose Grid applications from abstract application components that are mapped against available Grid services by the component framework at runtime. Copyright © 2006 John Wiley & Sons, Ltd. Rainer Schmidt 0003, Siegfried Benkner, Ivona Brandic, Gerhard Engelbrecht |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Quality of Service Negotiation for Commercial Medical Grid Services
Stuart E. Middleton, Mike Surridge, Siegfried Benkner, Gerhard Engelbrecht |
J. Grid Comput. | 3 |
| 2005 | QoS Support for Time-Critical Grid Workflow ApplicationsabstractTime critical grid applications as for example simulations for medical surgery or disaster recovery have special quality of service requirements. The Vienna Grid Environment, developed and evaluated in the context of the EU Project GEMSS, facilitates the provision of HPC applications as QoS-aware grid services by providing support for dynamic negotiation of various QoS guarantees like required execution time and price. In this paper, we extend the QoS mechanisms offered by the Vienna Grid Environment to workflow applications. We describe QoS extensions of the business process execution language and present a first prototype of a corresponding QoS-aware workflow engine which implements different strategies in order to bind the tasks of a workflow to adequate grid services subject to user-specified QoS constraints. We present different grid workflow planning approaches as well as first experimental results Ivona Brandic, Siegfried Benkner, Gerhard Engelbrecht, Rainer Schmidt 0003 |
e-Science | 2 |
| 2004 | Topic 4: Compilers for High Performance
Hans P. Zima, Siegfried Benkner, Michael F. P. O'Boyle, Beniamino Di Martino |
Euro-Par | 2 |
| 2004 | Compiling data-parallel programs for clusters of SMPsabstractAbstract Clusters of shared‐memory multiprocessors (SMPs) have become the most promising parallel computing platforms for scientific computing. However, SMP clusters significantly increase the complexity of user application development when using the low‐level application programming interfaces MPI and OpenMP, forcing users to deal with both distributed‐memory and shared‐memory parallelization details. In this paper we present extensions of High Performance Fortran (HPF) for SMP clusters which enable the compiler to adopt a hybrid parallelization strategy, efficiently combining distributed‐memory with shared‐memory parallelism. By means of a small set of new language features, the hierarchical structure of SMP clusters may be specified. This information is utilized by the compiler to derive inter‐node data mappings for controlling distributed‐memory parallelization across the nodes of a cluster and intra‐node data mappings for extracting shared‐memory parallelism within nodes. Additional mechanisms are proposed for specifying inter‐ and intra‐node data mappings explicitly, for controlling specific shared‐memory parallelization issues and for integrating OpenMP routines in HPF applications. The proposed features have been realized within the ADAPTOR and VFC compilers. The parallelization strategy for clusters of SMPs adopted by these compilers is discussed as well as a hybrid‐parallel execution model based on a combination of MPI and OpenMP. Experimental results indicate the effectiveness of the proposed features. Copyright © 2004 John Wiley & Sons, Ltd. Siegfried Benkner, Thomas Brandes |
Concurr. Comput. Pract. Exp. | 1 |
| 2003 | Performance of Java Web Services Implementations
Siegfried Benkner, Ivona Brandic, Aleksandar Dimitrov, Gerhard Engelbrecht, Rainer Schmidt 0003, Nikolay Terziev |
ICWS | 1 |
| 2002 | Efficient parallel programming on scalable shared memory systems with High Performance FortranabstractAbstract OpenMP offers a high‐level interface for parallel programming on scalable shared memory (SMP) architectures. It provides the user with simple work‐sharing directives while it relies on the compiler to generate parallel programs based on thread parallelism. However, the lack of language features for exploiting data locality often results in poor performance since the non‐uniform memory access times on scalable SMP machines cannot be neglected. High Performance Fortran (HPF), the de‐facto standard for data parallel programming, offers a rich set of data distribution directives in order to exploit data locality, but it has been mainly targeted towards distributed memory machines. In this paper we describe an optimized execution model for HPF programs on SMP machines that avails itself with mechanisms provided by OpenMP for work sharing and thread parallelism, while exploiting data locality based on user‐specified distribution directives. Data locality does not only ensure that most memory accesses are close to the executing threads and are therefore faster, but it also minimizes synchronization overheads, especially in the case of unstructured reductions. The proposed shared memory execution model for HPF relies on a small set of language extensions, which resemble the OpenMP work‐sharing features. These extensions, together with an optimized shared memory parallelization and execution model, have been implemented in the ADAPTOR HPF compilation system and experimental results verify the efficiency of the chosen approach. Copyright © 2002 John Wiley & Sons, Ltd. Siegfried Benkner, Thomas Brandes |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | High-performance numerical pricing methodsabstractAbstract The pricing of financial derivatives is an important field in finance and constitutes a major component of financial management applications. The uncertainty of future events often makes analytic approaches infeasible and, hence, time‐consuming numerical simulations are required. In the Aurora Financial Management System, pricing is performed on the basis of lattice representations of stochastic multidimensional scenario processes using the Monte Carlo simulation and Backward Induction methods, the latter allowing for the exploitation of shared‐memory parallelism. We present the parallelization of a Backward Induction numerical pricing kernel on a cluster of SMPs using HPF+, an extended version of High‐Performance Fortran. Based on language extensions for specifying a hierarchical mapping of data onto an SMP cluster, the compiler generates a hybrid‐parallel program combining distributed‐memory and shared‐memory parallelism. We outline the parallelization strategy adopted by the VFC compiler and present an experimental evaluation of the pricing kernel on an NEC SX‐5 vector supercomputer and a Linux SMP cluster, comparing a pure MPI version to a hybrid‐parallel MPI/OpenMP version. Copyright © 2002 John Wiley & Sons, Ltd. Hans Moritsch, Siegfried Benkner |
Concurr. Comput. Pract. Exp. | 2 |
| 2001 | High-Level Data Mapping for Clusters of SMPs
Siegfried Benkner, Thomas Brandes |
HIPS | 1 |
| 2001 | High-Level Data Mapping for Clusters of SMPs
Siegfried Benkner, Thomas Brandes |
IPDPS | 1 |
| 2000 | Exploiting Data Locality on Scalable Shared Memory Machines with Data Parallel Programs
Siegfried Benkner, Thomas Brandes |
Euro-Par | 1 |
| 2000 | Optimizing Irregular HPF Applications Using HalosabstractThis paper presents extensions of High Performance Fortran (HPF) for specifying non-local access patterns of distributed arrays, called halos, and for controlling communication associated with these non-local accesses. Using these features, crucial optimization techniques required for an efficient parallelization of irregular applications may be applied. The information provided by halos is utilized by the compiler and the runtime system for optimizing the management of distributed arrays and the computation of communication schedules. High-level communication primitives for halos enable the programmer to avoid redundant communication, to reuse communication schedules, and to hide communication overheads by overlapping communication with computation. Performance results of a kernel from a crash simulation application on the NEC Cenju-4, the IBM SP2, and on the NEC SX-4 demonstrate that by using the proposed extensions, a performance close to handwritten message-passing codes can be achieved for irregular problems. Copyright © 2000 John Wiley & Sons, Ltd. Siegfried Benkner |
Concurr. Pract. Exp. | 1 |
| 1999 | The HPF+ Project: Supporting HPF for Advanced Industrial Applications
Siegfried Benkner, Guy Lonsdale, Hans P. Zima |
Euro-Par | 1 |
| 1999 | HPF+: High Performance Fortran for advanced scientific and engineering applications
Siegfried Benkner |
Future Gener. Comput. Syst. | 1 |
| 1999 | Compiling High Performance Fortran for distributed-memory architectures
Siegfried Benkner, Hans P. Zima |
Parallel Comput. | 1 |
| 1998 | High-level Management of Communication Schedules in HPF-like LanguagesabstractThe goal of High Performance Fortran (HPF) is to "address the problems of writing data parallel programs where the distribution of data affects performance", Siegfried Benkner, Piyush Mehrotra, John Van Rosendale, Hans P. Zima |
International Conference on Supercomputing | 1 |
| 1995 | Handling block-cyclic distributed arrays in Vienna Fortran 90
Siegfried Benkner |
PACT | 1 |
| 1994 | Processing Array Statements and Procedure Interfaces in the PREPARE High Performance Fortran Compiler
Siegfried Benkner, Peter Brezany, Hans P. Zima |
CC | 1 |