Clemens Grelck

dblp:54/1303 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
4since 2021 · last 2023
0000-0003-3003-1388ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 6 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 The TeamPlay Project: Analysing and Optimising Time, Energy, and Security for Cyber-Physical Systems
abstract
Non-functional properties, such as energy, time, and security (ETS) are becoming increasingly important in Cyber-Physical Systems (CPS) programming. This article describes TeamPlay, a research project funded under the EU Horizon 2020 programme between January 2018 and June 2021. TeamPlay aimed to provide the system designer with a toolchain for developing embedded applications where ETS properties are first-class citizens, allowing the developer to reflect directly on energy, time and security properties at the source code level. In this paper we give an overview of the TeamPlay methodology, introduce the challenges and solutions of our approach and summarise the results achieved. Overall, applying our TeamPlay methodology led to an improvement of up to 18% performance and 52% energy usage over traditional approaches.
Benjamin Rouxel, Christopher Brown 0002, Emad Samuel Malki Ebeid, Kerstin Eder, Heiko Falk, Clemens Grelck, Jesper Holst, Shashank Jadhav, Yoann Marquer, Marcos Martinez de Alejandro, Kris Nikov, Ali Sahafi, Ulrik Pagh Schultz Lundquist, Adam Seewald, Vangelis Vassalos, Simon Wegener, Olivier Zendra
DATE6
2023 Change of Plans: Optimizing for Power, Reliability and Timeliness for Cost-Conscious Real-Time Systems
abstract
The large-scale adoption of embedded devices naturally sees them deployed under widely varying constraints matching an equally wide number of environments and applications. Such devices may be under hard timing constraints as a result of their domain (real-time), but other constraints, such as reliability and power, are becoming increasingly relevant. In particular, battery-powered devices are limited in runtime due to their long-running average power (amortized power) consumption. We propose a real-time scheduling algorithm optimizing for reliability under an amortized energy budget. Weakly-hard real-time tasks can tolerate patterns of failures and successes, and we leverage this property as the driver of reliability in a constrained environment. Reliability against transient faults is increased by dynamically switching the protected subset of the task set based on the observed history of task success/fails. By performing such a change of plans only for tasks in need, we can cost-effectively boost reliability while lowering energy consumption. Our method optimizes reliability for a given deadline and energy budget. We explore the behavior of our method under a wide range of energy and time constrained environments, illustrating the trade-off between amortized power, timeliness and reliability, enabling application designers to make cost-conscious decisions.
Lukas Miedema, Clemens Grelck
DSD2
2022 Strategy Switching: Smart Fault-Tolerance for Weakly-Hard Resource-Constrained Real-Time Applications
Lukas Miedema, Clemens Grelck
SEFM2
2021 YASMIN: a real-time middleware for COTS heterogeneous platforms
abstract
Commercial-off-the-shelf (COTS) heterogeneous platforms provide immense computational power, but are difficult to program and to correctly use when real-time requirements come into play: A sound configuration of the operating system scheduler is needed, and a suitable mapping of tasks to computing units must be determined. Flawed designs lead to sub-optimal system configurations and, thus, to wasted resources or even to deadline misses and system failures.
Benjamin Rouxel, Sebastian Altmeyer, Clemens Grelck
Middleware3
2020 Q-learning for Statically Scheduling DAGs
abstract
Data parallel frameworks (e.g. Hive, Spark or Tez) can be used to execute complex data analyses consisting of many dependent tasks represented by a Directed Acylical Graph (DAG). Minimising the job completion time (i.e. makespan) is still an open problem for large graphs.We propose a novel deep Q-learning (DQN) approach to statically scheduling DAGs and minimising the makespan. Our approach learns to schedule DAGs from scratch instead of learning how to imitate some heuristic. We show that our current approach learns fast and steadily. Furthermore, our approach can schedule DAGs almost 15 times faster than a Forward List Scheduling (FLS) heuristic.
Julius Roeder, Benjamin Rouxel, Clemens Grelck
IEEE BigData3
2020 Towards Energy-, Time- and Security-Aware Multi-core Coordination
Julius Roeder, Benjamin Rouxel, Sebastian Altmeyer, Clemens Grelck
COORDINATION4
2020 PReGO: a generative methodology for satisfying real-time requirements on COTS-based systems: definition and experience report
abstract
Satisfying real-time requirements in cyber-physical systems is challenging as timing behaviour depends on the application software, the embedded hardware, as well as the execution environment. This challenge is exacerbated as real-world, industrial systems often use unpredictable hardware and software libraries or operating systems with timing hazards and proprietary device drivers. All these issues limit or entirely prevent the application of established real-time analysis techniques.
Benjamin Rouxel, Ulrik Pagh Schultz Lundquist, Benny Akesson, Jesper Holst, Ole Jørgensen 0001, Clemens Grelck
GPCE6
2020 Programming languages for data-Intensive HPC applications: A systematic mapping study
Vasco Amaral 0001, Beatriz Norberto, Miguel Goulão, Marco Aldinucci, Siegfried Benkner, Andrea Bracciali, Paulo Carreira 0001, Edgars Celms, Luís Correia 0001, Clemens Grelck, Helen D. Karatza, Christoph W. Kessler, Peter Kilpatrick, Hugo F. M. C. Martiniano, Ilias Mavridis, Sabri Pllana, Ana Respício, José Simão, Luís Veiga, Ari Visa
Parallel Comput.10
2019 SAC Goes Cluster: Fully Implicit Distributed Computing
abstract
SAC (Single Assignment C) is a purely functional, data-parallel array programming language that predominantly targets compute-intensive applications. Thus, clusters of workstations, or distributed memory architectures in general, form highly relevant compilation targets. Notwithstanding, SAC as of today only supports shared-memory architectures, graphics accelerators and heterogeneous combinations thereof. In our current work we aim at closing this gap. At the same time, we are determined to uphold SAC's promise of entirely compiler-directed exploitation of concurrency, no matter what the target architecture is. Distributed memory architectures are going to make this promise a particular challenge. Despite SAC's functional semantics, it is generally far from straightforward to infer exact communication patterns from architecture-agnostic code. Therefore, we intend to capitalise on recent advances in network technology, namely the closing of the gap between memory bandwidth and network bandwidth. We aim at a solution based on a custom-designed software distributed shared memory (S-DSM) and large per-node software-managed cache memories. To this effect the functional nature of SAC with its write-once/read-only arrays provides a strategic advantage that we thoroughly exploit. Throughout the paper we further motivate our approach, sketch out our implementation strategy, show preliminary results and discuss the pros and cons of our approach.
Thomas Macht, Clemens Grelck
IPDPS2
2014 SaC/C formulations of the all-pairs N-body problem and their performance on SMPs and GPGPUs
abstract
SUMMARY This paper describes our experience in implementing the classical N‐body algorithm in SaC and analysing the runtime performance achieved on three different machines: a dual‐processor 8‐core Dell PowerEdge 2950 (a Beowulf cluster node, the reference machine), a quad‐core hyper‐threaded Intel Core‐i7 based system equipped with an NVidia GTX‐480 graphics accelerator and an Oracle Sparc T4‐4 server with a total of 256 hardware threads. We contrast our findings with those resulting from the reference C code and a few variants of it that employ OpenMP pragmas as well as explicit vectorisation. Our experiments demonstrate that the SaC implementation successfully combines a high level of abstraction, very close to the mathematical specification, with very competitive runtimes. In fact, SaC matches or outperforms the hand‐vectorised and hand‐parallelised C codes on all three systems under investigation without the need for any source code modification. Furthermore, only SaC is able to effectively harness the advanced compute power of the graphics accelerator, again by mere recompilation of the same source code. Our results illustrate the benefits that SaC provides to application programmers in terms of coding productivity, source code, and performance portability among different machine architectures, as well as long‐term maintainability in evolving hardware environments. Copyright © 2013 John Wiley & Sons, Ltd.
Artjoms Sinkarovs, Sven-Bodo Scholz, Robert Bernecky, Roeland Douma, Clemens Grelck
Concurr. Comput. Pract. Exp.5
2012 Distributed S-Net: Cluster and Grid Computing without the Hassle
abstract
S-Net is a declarative coordination language and component technology primarily aimed at modern multi-core/many-core chip architectures. It builds on the concept of stream processing to structure dynamically evolving networks of communicating asynchronous components, which themselves are implemented using a conventional language suitable for the application domain. We present the design and implementation of Distributed S-Net, a conservative extension of S-Net aimed at distributed memory architectures ranging from many-core chip architectures with hierarchical memory organisations to more traditional clusters of workstations, supercomputers and grids. Three case studies illustrate how to use Distributed S-Net to implement different models of parallel execution. Runtimes obtained on a workstation cluster demonstrate how Distributed S-Net allows programmers with little or no background in parallel programming to make effective use of distributed memory architectures with minimal programming effort.
Clemens Grelck, Jukka Julku, Frank Penczek
CCGRID1
2012 Asynchronous adaptive optimisation for generic data-parallel array programming
abstract
SUMMARY Programming productivity very much depends on the availability of basic building blocks that can be reused for a wide range of application scenarios and the ability to define rich abstraction hierarchies. Driven by the aim for increased reuse, such basic building blocks tend to become more and more generic in their specification; structural as well as behavioural properties are turned into parameters that are passed on to lower layers of abstraction where eventually a differentiation is being made. In the context of array programming, such properties are typically array ranks (number of axes/dimensions) and array shapes (number of elements along each axis/dimension). This allows for abstract definitions of operations such as element‐wise additions, concatenations, rotations, and so on, which jointly enable a very high‐level compositional style of programming, similar to, for instance, MATLAB. However, such a generic programming style generally comes at a price in terms of runtime overheads when compared against tailor‐made low‐level implementations. Additional layers of abstraction as well as the lack of hard‐coded structural properties often inhibits optimisations that are obvious otherwise. Although complex static compiler analyses and transformations such as partial evaluations can ameliorate the situation to quite some extent, there are cases, where the required level of information is not available until runtime. In this paper, we propose to shift part of the optimisation process into the runtime of applications. Triggered by some runtime observation, the compiler asynchronously applies partial evaluation techniques to frequently used program parts and dynamically replaces initial program fragments by more specialised ones through dynamic re‐linking. In contrast to many existing approaches, we suggest this optimisation to be done in a rather non‐intrusive, decoupled way. We use a full‐fledged compiler that is run on a separate core. This measure enables us to run the compiler on its highest optimisation‐level, which requires non‐negligible compilation times for our optimisations. We use the compiler's type system to identify the potential dynamic optimisations. And we use the host language's module system as a facilitator for the dynamic code modifications. We present the architecture and implementation of an adaptive compilation framework for Single Assignment C, a data‐parallel array programming language. Single Assignment C advocates shape‐generic and rank‐generic programming with arrays. A sophisticated, highly optimising compiler technology nevertheless achieves competitive runtime performance. We demonstrate the suitability of our approach to achieve consistently high performance independent of the static availability of array properties by means of several experiments based on a highly generic formulation of rank‐invariant convolution as a case study. Copyright © 2011 John Wiley & Sons, Ltd.
Clemens Grelck, Tim van Deurzen, Stephan Herhut, Sven-Bodo Scholz
Concurr. Comput. Pract. Exp.1
2010 Cluster Computing as an Assembly Process: Coordination with S-Net
abstract
This poster will present a coordination language for distributed computing and will discuss its application to cluster computing. It will introduce a programming technique of cluster computing whereby application components are completely dissociated from the communication/coordination infrastructure (unlike MPI-style message passing), and there is no shared memory either, whether virtual or physical (unlike Open-MP). Cluster computing is thus presented as something that happens as late as the assembly stage: components are integrated into an application using a new form of network glue: Single-Input, Single-Output (SISO) asynchronous, no deterministic coordination.
Clemens Grelck, Jukka Julku, Frank Penczek, Alexander V. Shafarenko
CCGRID1
2007 Coordinating Data Parallel SAC Programs with S-Net
abstract
We propose a two-layered approach for exploiting different forms of concurrency in complex systems: we specify computational components in our functional array language SAC, which exploits data parallel properties of array processing code. The declarative stream processing language S-Net is used to orchestrate the collaborative behaviour of these components in a streaming network. We illustrate our approach by a hybrid implementation of a sudoku puzzle solver as a representative for more complex search problems.
Clemens Grelck, Sven-Bodo Scholz, Alexander V. Shafarenko
IPDPS1
2006 Merging compositions of array skeletons in SaC
Clemens Grelck, Sven-Bodo Scholz
Parallel Comput.1
2005 Shared memory multiprocessor support for functional array processing in SAC
abstract
Classical application domains of parallel computing are dominated by processing large arrays of numerical data. Whereas most functional languages focus on lists and trees rather than on arrays, S A C is tailor-made in design and in implementation for efficient high-level array processing. Advanced compiler optimizations yield performance levels that are often competitive with low-level imperative implementations. Based on S A C, we develop compilation techniques and runtime system support for the compiler-directed parallel execution of high-level functional array processing code on shared memory architectures. Competitive sequential performance gives us the opportunity to exploit the conceptual advantages of the functional paradigm for achieving real performance gains with respect to existing imperative implementations, not only in comparison with uniprocessor runtimes. While the design of S A C facilitates parallelization, the particular challenge of high sequential performance is that realization of satisfying speedups through parallelization becomes substantially more difficult. We present an initial compilation scheme and multi-threaded execution model, which we step-wise refine to reduce organizational overhead and to improve parallel performance. We close with a detailed analysis of the impact of certain design decisions on runtime performance, based on a series of experiments.
Clemens Grelck
J. Funct. Program.1
2000 HPF vs. SAC - A Case Study (Research Note)
Clemens Grelck, Sven-Bodo Scholz
Euro-Par1