Erwin Laure

dblp:15/4220 · DBLP profile ↗
← Back
50ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-9901-9857ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 8 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Performance Model of In-Situ Techniques
Yi Ju, Nicolas Vidal 0003, Adalberto Perez, Ana Gainaru, Frédéric Suter, Stefano Markidis, Philipp Schlatter, Scott Klasky, Erwin Laure
PDP9
2024 Dynamic Resource Management for In-Situ Techniques Using MPI-Sessions
Yi Ju, Dominik Huber, Adalberto Perez, Philipp Ulbl, Stefano Markidis, Philipp Schlatter, Martin Schulz 0001, Martin Schreiber 0001, Erwin Laure
EuroMPI9
2023 In-Situ Techniques on GPU-Accelerated Data-Intensive Applications
abstract
The computational power of High-Performance Computing (HPC) systems is constantly increasing, however, their input/output (IO) performance grows relatively slowly, and their storage capacity is also limited. This unbalance presents significant challenges for applications such as Molecular Dynamics (MD) and Computational Fluid Dynamics (CFD), which generate massive amounts of data for further visualization or analysis. At the same time, checkpointing is crucial for long runs on HPC clusters, due to limited walltimes and/or failures of system components, and typically requires the storage of large amount of data. Thus, restricted IO performance and storage capacity can lead to bottlenecks for the performance of full application workflows (as compared to computational kernels without IO). In-situ techniques, where data is further processed while still in memory rather to write it out over the I/O subsystem, can help to tackle these problems. In contrast to traditional post-processing methods, in-situ techniques can reduce or avoid the need to write or read data via the IO subsystem. They offer a promising approach for applications aiming to leverage the full power of large scale HPC systems. In-situ techniques can also be applied to hybrid computational nodes on HPC systems consisting of graphics processing units (GPUs) and central processing units (CPUs). On one node, the GPUs would have significant performance advantages over the CPUs. Therefore, current approaches for GPU-accelerated applications often focus on maximizing GPU usage, leaving CPUs underutilized. In-situ tasks using CPUs to perform data analysis or preprocess data concurrently to the running simulation, offer a possibility to improve this underutilization.
Yi Ju, Mingshuai Li, Adalberto Perez, Laura Bellentani, Niclas Jansson, Stefano Markidis, Philipp Schlatter, Erwin Laure
e-Science8
2023 Beyond the Fourth Paradigm - the Rise of AI
abstract
Thanks to the availability of huge amounts of data and improved computational resources, AI methods are gaining importance in scientific workflows, from image recognition and natural language processing to materials science. In many domains the usage of AI is under active investigation and first results show a tremendous potential, suggesting that AI will have significant impact way beyond the currently dominating examples of image and language processing.
Andreas Marek, Markus Rampp, Klaus Reuter, Erwin Laure
e-Science4
2023 Exploring the Ultimate Regime of Turbulent Rayleigh-Bénard Convection Through Unprecedented Spectral-Element Simulations
abstract
We detail our developments in the high-fidelity spectral-element code Neko that are essential for unprecedented large-scale direct numerical simulations of fully developed turbulence. Major innovations are modular multi-backend design enabling performance portability across a wide range of GPUs and CPUs, a GPU-optimized preconditioner with task overlapping for the pressure-Poisson equation and in-situ data compression. We carry out initial runs of Rayleigh-Bénard Convection (RBC) at extreme scale on the LUMI and Leonardo supercomputers. We show how Neko is able to strongly scale to 16,384 GPUs and obtain results that are not possible without careful consideration and optimization of the entire simulation workflow. These developments in Neko will help resolving the long-standing question regarding the ultimate regime in RBC.
Niclas Jansson, Martin Karp, Adalberto Perez, Timofey Mukha, Yi Ju, Jiahui Liu 0006, Szilárd Páll, Erwin Laure, Tino Weinkauf, Jörg Schumacher, Philipp Schlatter, Stefano Markidis
SC8
2022 A Performance Evaluation of Adaptive MPI for a Particle-In-Cell Code
abstract
In the quest for extreme-scale supercomputers, the High Performance Computing (HPC) community has developed many resources (programming paradigms, architectures, method-ologies, numerical methods) to face the multiple challenges along the way. One of those resources are task-based parallel program-ming tools. The availability of mature programming models, pro-gramming languages, and runtime systems that use task-based parallelism represent a favorable ecosystem. The fundamental premise of these tools is their ability to naturally cope with dynamically changing execution conditions, i.e. adaptivity. In this paper, we explore Adaptive MPI, a parallel-object framework, as a mechanism to provide, among other features, automatic and dynamic load balancing for a particle-in-cell application. We ported a pre-existing MPI application on the Adaptive MPI infrastructure and highlight the changes required to the code. Our experimental results show Adaptive MPI has a minimum overhead, maintains the scalability of the original code, and it is able to alleviate an artificially-introduced load imbalance.
Christian Asch, Diego Jiménez, Markus Rampp, Erwin Laure, Esteban Meneses
CLUSTER4
2022 Understanding the Impact of Synchronous, Asynchronous, and Hybrid In-Situ Techniques in Computational Fluid Dynamics Applications
abstract
High-Performance Computing (HPC) systems provide input/output (IO) performance growing relatively slowly compared to peak computational performance and have limited storage capacity. Computational Fluid Dynamics (CFD) applications aiming to leverage the full power of Exascale HPC systems, such as the solver Nek5000, will generate massive data for further processing. These data need to be efficiently stored via the IO subsystem. However, limited IO performance and storage capacity may result in performance, and thus scientific discovery, bottlenecks. In comparison to traditional post-processing methods, in-situ techniques can reduce or avoid writing and reading the data through the IO subsystem, promising to be a solution to these problems. In this paper, we study the performance and resource usage of three in-situ use cases: data compression, image generation, and uncertainty quantification. We furthermore analyze three approaches when these in-situ tasks and the simulation are executed synchronously, asynchronously, or in a hybrid manner. In-situ compression can be used to reduce the IO time and storage requirements while maintaining data accuracy. Furthermore, in-situ visualization and analysis can save Terabytes of data from being routed through the IO subsystem to storage. However, the overall efficiency is crucially dependent on the characteristics of both, the in-situ task and the simulation. In some cases, the overhead introduced by the in-situ tasks can be substantial. Therefore, it is essential to choose the proper in-situ approach, synchronous, asynchronous, or hybrid, to minimize overhead and maximize the benefits of concurrent execution.
Yi Ju, Adalberto Perez, Stefano Markidis, Philipp Schlatter, Erwin Laure
e-Science5
2022 Strong Scaling of OpenACC enabled Nek5000 on several GPU based HPC systems
abstract
We present new results on the strong parallel scaling for the OpenACC-accelerated implementation of the high-order spectral element fluid dynamics solver Nek5000. The test case considered consists of a direct numerical simulation of fully-developed turbulent flow in a straight pipe, at two different Reynolds numbers Reτ = 360 and Reτ = 550, based on friction velocity and pipe radius. The strong scaling is tested on several GPU-enabled HPC systems, including the Swiss Piz Daint system, TACC’s Longhorn, Jülich’s JUWELS Booster, and Berzelius in Sweden. The performance results show that speed-up between 3-5 can be achieved using the GPU accelerated version compared with the CPU version on these different systems. The run-time for 20 timesteps reduces from 43.5 to 13.2 seconds with increasing the number of GPUs from 64 to 512 for Reτ = 550 case on JUWELS Booster system. This illustrates the GPU accelerated version the potential for high throughput. At the same time, the strong scaling limit is significantly larger for GPUs, at about 2000 − 5000 elements per rank; compared to about 50 − 100 for a CPU-rank.
Jonathan Vincent, Martin Karp, Adam Peplinski, Niclas Jansson, Artur Podobas, Andreas Jocksch, Fazle Hussain, Stefano Markidis, Matts Karlsson, Dirk Pleiter, Erwin Laure, Philipp Schlatter
HPC Asia13
2022 Exploiting Reduced Precision for GPU-based Time Series Mining
abstract
The mining of multi-dimensional time series is a crucial step in gaining insights into data obtained from physical systems and from monitoring infrastructures. A widely accepted approach for this challenge is the matrix profile, which, however, is computationally very expensive. It relies on calculating large correlation matrices coupled with sort operations across all dimensions of the data, as well as on performing inclusive scans. All of these steps are inherently data parallel and can, therefore, benefit from execution on GPUs, and even more so from horizontal scaling on multiple GPUs. In addition, the nature of the matrix profile calculation allows the exploitation of reduced precision on GPUs. This offers further improvements to enable the analysis of ever growing data sets in real-world scenarios. Based on these motivations, we introduce the first parallel algorithm for multi-dimensional matrix profile on multiple GPUs exploiting reduced precision modes and provide a highly opti-mized implementation using novel optimization techniques. On one NVIDIA A100 GPU, our implementation achieves a 54x performance improvement in comparison to an optimized single-node execution on a state-of-the-art CPU-based implementation relying on double-precision computation and an additional factor of 1.4x when switching to reduced precision while maintaining sufficient accuracy. We study the accuracy and performance trade-offs for our proposed algorithm in detail and present synthetic and real-world case studies to demonstrate how the reduced precision improves the performance, while accomplishing sufficiently accurate results.
Yi Ju, Amir Raoofy, Dai Yang, Erwin Laure, Martin Schulz 0001
IPDPS4
2022 In situ visualization of large-scale turbulence simulations in Nek5000 with ParaView Catalyst
abstract
Abstract In situ visualization on high-performance computing systems allows us to analyze simulation results that would otherwise be impossible, given the size of the simulation data sets and offline post-processing execution time. We develop an in situ adaptor for Paraview Catalyst and Nek5000, a massively parallel Fortran and C code for computational fluid dynamics. We perform a strong scalability test up to 2048 cores on KTH’s Beskow Cray XC40 supercomputer and assess in situ visualization’s impact on the Nek5000 performance. In our study case, a high-fidelity simulation of turbulent flow, we observe that in situ operations significantly limit the strong scalability of the code, reducing the relative parallel efficiency to only $$\approx 21\%$$ ≈ 21 % on 2048 cores (the relative efficiency of Nek5000 without in situ operations is $$\approx 99\%$$ ≈ 99 % ). Through profiling with Arm MAP, we identified a bottleneck in the image composition step (that uses the Radix-kr algorithm) where a majority of the time is spent on MPI communication. We also identified an imbalance of in situ processing time between rank 0 and all other ranks. In our case, better scaling and load-balancing in the parallel image composition would considerably improve the performance of Nek5000 with in situ capabilities. In general, the result of this study highlights the technical challenges posed by the integration of high-performance simulation codes and data-analysis libraries and their practical use in complex cases, even when efficient algorithms already exist for a certain application scenario.
Marco Atzori, Wiebke Köpp, Steven W. D. Chien, Daniele Massaro, Fermín Mallor, Adam Peplinski, Mohamad Rezaei, Niclas Jansson, Stefano Markidis, Ricardo Vinuesa, Erwin Laure, Philipp Schlatter, Tino Weinkauf
J. Supercomput.11
2019 Hybrid Resource Management for HPC and Data Intensive Workloads
abstract
High Performance Computing (HPC) and Data Intensive (DI) workloads have been executed on separate clusters using different tools for resource and application management. With increasing convergence, where modern applications are composed of both types of jobs in complex workflows, this separation becomes a growing overhead and the need for a common platform increases. Executing both workload classes on the same clusters not only enables hybrid workflows, but can also increase system efficiency, as available hardware often is not fully utilized by applications. While HPC systems are typically managed in a coarse grained fashion, with exclusive resource allocations, DI systems employ a finer grained regime, enabling dynamic allocation and control based on application needs. On the path to full convergence, a useful and less intrusive step is a hybrid resource management system allowing the execution of DI applications on top of standard HPC scheduling systems. In this paper we present the architecture of a hybrid system enabling dual-level scheduling for DI jobs in HPC infrastructures. Our system takes advantage of real-time resource profiling to efficiently co-schedule HPC and DI applications. The architecture is easily extensible to current and new types of distributed applications, allowing efficient combination of hybrid workloads on HPC resources with increased job throughput and higher overall resource utilization. The implementation is based on the Slurm and Mesos resource managers for HPC and DI jobs. Experimental evaluations in a real cluster based on a set of representative HPC and DI applications demonstrate that our hybrid architecture improves resource utilization by 20%, with 12% decrease on queue makespan while still meeting all deadlines for HPC jobs.
Abel Souza, Mohamad Rezaei, Erwin Laure, Johan Tordsson
CCGRID3
2019 The Future of Swedish e-Science: SeRC 2.0
abstract
Since 2010, the Swedish e-Science Research Centre (SeRC) is funding and coordinating e-Science activities in a broad spectrum of scientific disciplines. After an initial 5-year phase that produced outstanding results, SeRC is increasingly focusing on fostering interactions between disciplines and has created so-called Multidisciplinary Collaborative Programs (MCPs). In these programs, domain researchers collaborate with e-Science methods and tool developers and e-Infrastructure providers. In this paper we give an overview of the initial phase of SeRC and present the new programs that started operating in 2019.
Erwin Laure, Olivia Eriksson, Erik Lindahl, Dan S. Henningson
eScience1
2019 OpenACC acceleration for the PN-PN-2 algorithm in Nek5000
Evelyn Otero, Misun Min, Paul F. Fischer, Philipp Schlatter, Erwin Laure
J. Parallel Distributed Comput.6
2019 SAGE: Percipient Storage for Exascale Data Centric Computing
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Stefano Markidis, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Dirk Pleiter, Shaun De Witt
Parallel Comput.8
2018 The SAGE project: a storage centric approach for exascale computing: invited paper
abstract
SAGE (Percipient StorAGe for Exascale Data Centric Computing) is a European Commission funded project towards the era of Exascale computing. Its goal is to design and implement a Big Data/Extreme Computing (BDEC) capable infrastructure with associated software stack. The SAGE system follows a storage centric approach as it is capable of storing and processing large data volumes at the Exascale regime.
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Steven W. D. Chien, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Shaun De Witt, Dirk Pleiter, Stefano Markidis
CF8
2018 Characterizing the performance benefit of hybrid memory system for HPC applications
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Jeffrey S. Vetter, Pietro Cicotti, Erwin Laure, Stefano Markidis
Parallel Comput.6
2018 MPI windows on storage for HPC applications
Sergio Rivas-Gomez, Roberto Gioiosa, Ivy Bo Peng, Gokcen Kestor, Sai Narasimhamurthy, Erwin Laure, Stefano Markidis
Parallel Comput.6
2018 A taxonomy of task-based parallel programming technologies for high-performance computing
abstract
Task-based programming models for shared memory—such as Cilk Plus and OpenMP 3—are well established and documented. However, with the increase in parallel, many-core, and heterogeneous systems, a number of research-driven projects have developed more diversified task-based support, employing various programming and runtime features. Unfortunately, despite the fact that dozens of different task-based systems exist today and are actively used for parallel and high-performance computing (HPC), no comprehensive overview or classification of task-based technologies for HPC exists. In this paper, we provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms. We demonstrate the usefulness of our taxonomy by classifying state-of-the-art task-based environments in use today.
Peter Thoman, Kiril Dichev, Thomas Heller, Roman Iakymchuk, Xavier Aguilar, Khalid Hasanov, Philipp Gschwandtner, Pierre Lemarinier, Stefano Markidis, Herbert Jordan, Thomas Fahringer, Kostas Katrinis, Erwin Laure, Dimitrios S. Nikolopoulos
J. Supercomput.13
2017 Extending Message Passing Interface Windows to Storage
abstract
This paper presents an extension to MPI supporting the one-sided communication model and window allocations in storage. Our design transparently integrates with the current MPI implementations, enabling applications to target MPI windows in storage, memory or both simultaneously, without major modifications. Initial performance results demonstrate that the presented MPI window extension could potentially be helpful for a wide-range of use-cases and with low-overhead.
Sergio Rivas-Gomez, Stefano Markidis, Ivy Bo Peng, Erwin Laure, Gokcen Kestor, Roberto Gioiosa
CCGrid4
2017 Preparing HPC Applications for the Exascale Era: A Decoupling Strategy
abstract
Production-quality parallel applications are often a mixture of diverse operations, such as computation- and communication-intensive, regular and irregular, tightly coupled and loosely linked operations. In conventional construction of parallel applications, each process performs all the operations, which might result inefficient and seriously limit scalability, especially at large scale. We propose a decoupling strategy to improve the scalability of applications running on large-scale systems. Our strategy separates application operations onto groups of processes and enables a dataflow processing paradigm among the groups. This mechanism is effective in reducing the impact of load imbalance and increases the parallel efficiency by pipelining multiple operations. We provide a proof-of-concept implementation using MPI, the de-facto programming system on current supercomputers. We demonstrate the effectiveness of this strategy by decoupling the reduce, particle communication, halo exchange and I/O operations in a set of scientific and data-analytics applications. A performance evaluation on 8,192 processes of a Cray XC40 supercomputer shows that the proposed approach can achieve up to 4x performance improvement.
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Erwin Laure, Stefano Markidis
ICPP4
2017 RTHMS: a tool for data placement on hybrid memory system
abstract
Traditional scientific and emerging data analytics applications require fast, power-efficient, large, and persistent memories. Combining all these characteristics within a single memory technology is expensive and hence future supercomputers will feature different memory technologies side-by-side. However, it is a complex task to program hybrid-memory systems and to identify the best object-to-memory mapping. We envision that programmers will probably resort to use default configurations that only require minimal interventions on the application code or system settings. In this work, we argue that intelligent, fine-grained data placement can achieve higher performance than default setups.
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Pietro Cicotti, Erwin Laure, Stefano Markidis
ISMM5
2017 E-Science technologies in a workflow for personalized medicine using cancer screening as a case study
abstract
OBJECTIVE: We provide an e-Science perspective on the workflow from risk factor discovery and classification of disease to evaluation of personalized intervention programs. As case studies, we use personalized prostate and breast cancer screenings. MATERIALS AND METHODS: We describe an e-Science initiative in Sweden, e-Science for Cancer Prevention and Control (eCPC), which supports biomarker discovery and offers decision support for personalized intervention strategies. The generic eCPC contribution is a workflow with 4 nodes applied iteratively, and the concept of e-Science signifies systematic use of tools from the mathematical, statistical, data, and computer sciences. RESULTS: The eCPC workflow is illustrated through 2 case studies. For prostate cancer, an in-house personalized screening tool, the Stockholm-3 model (S3M), is presented as an alternative to prostate-specific antigen testing alone. S3M is evaluated in a trial setting and plans for rollout in the population are discussed. For breast cancer, new biomarkers based on breast density and molecular profiles are developed and the US multicenter Women Informed to Screen Depending on Measures (WISDOM) trial is referred to for evaluation. While current eCPC data management uses a traditional data warehouse model, we discuss eCPC-developed features of a coherent data integration platform. DISCUSSION AND CONCLUSION: E-Science tools are a key part of an evidence-based process for personalized medicine. This paper provides a structured workflow from data and models to evaluation of new personalized intervention strategies. The importance of multidisciplinary collaboration is emphasized. Importantly, the generic concepts of the suggested eCPC workflow are transferrable to other disease domains, although each disease will require tailored solutions.
Ola Spjuth, Andreas Rosenblad, Mark A. Clements, Keith Humphreys, Emma Ivansson, Jim Dowling, Martin Eklund, Alexandra Jauhiainen, Kamila Czene, Henrik Gronberg, Pär Sparén, Fredrik Wiklund, Abbas Cheddad, þorgerður Pálsdóttir, Mattias Rantalainen, Linda Abrahamsson, Erwin Laure, Jan-Eric Litton, Juni Palmgren
J. Am. Medical Informatics Assoc.17
2016 A parallel microsimulation package for modelling cancer screening policies
abstract
Microsimulation with stochastic life histories is an important tool in the development of public policies. In this article, we use microsimulation to evaluate policies for prostate cancer testing. We implemented the microsimulations as an R package, with pre- and post-processing in R and with the simulations written in C++. Calibrating a microsimulation model with a large population can be computationally expensive. To address this issue, we investigated four forms of parallelism: (i) shared memory parallelism using R; (ii) shared memory parallelism using OpenMP at the C++ level; (iii) distributed memory parallelism using R; and (iv) a hybrid shared/distributed memory parallelism using OpenMP at the C++ level and MPI at the R level. The close coupling between R and C++ offered advantages for ease of software dissemination and the use of high-level R parallelisation methods. However, this combination brought challenges when trying to use shared memory parallelism at the C++ level: the performance gained by hybrid OpenMP/MPI came at the cost of significant re-factoring of the existing code. As a case study, we implemented a prostate cancer model in the microsimulation package. We used this model to investigate whether prostate cancer testing with specific re-testing protocols would reduce harms and maintain any mortality benefit from prostate-specific antigen testing. We showed that four-yearly testing would have a comparable effectiveness and a marked decrease in costs compared with two-yearly testing and current testing. In summary, we developed a microsimulation package in R and assessed the cost-effectiveness of prostate cancer testing. We were able to scale up the microsimulations using a combination of R and C++, however care was required when using shared memory parallelism at the C++ level.
Andreas Rosenblad, Niten Olofsson, Erwin Laure, Mark A. Clements
eScience3
2016 Nekbone performance on GPUs with OpenACC and CUDA Fortran implementations
Stefano Markidis, Erwin Laure, Matthew Otten, Paul F. Fischer, Misun Min
J. Supercomput.3
2015 On the Application Task Granularity and the Interplay with the Scheduling Overhead in Many-Core Shared Memory Systems
abstract
Task-based programming models are considered one of the most promising programming model approaches for exascale supercomputers because of their ability to dynamically react to changing conditions and reassign work to processing elements. One question, however, remains unsolved: what should the task granularity of task-based applications be? Fine-grained tasks offer more opportunities to balance the system and generally result in higher system utilization. However, they also induce in large scheduling overhead. The impact of scheduling overhead on coarse-grained tasks is lower, but large systems may result imbalanced and underutilized. In this work we propose a methodology to analyze the interplay between application task granularity and scheduling overhead. Our methodology is based on three main points: 1) a novel task algorithm that analyzes an application directed acyclic graph (DAG) and aggregates tasks, 2) a fast and precise emulator to analyze the application behavior on systems with up to 1,024 cores, 3) a comprehensive sensitivity analysis of application performance and scheduling overhead breakdown. Our results show that there is an optimal task granularity between 1.2×104and 10×104cycles for the representative schedulers. Moreover, our analysis indicates that a suitable scheduler for exascale task-based applications should employ a best-effort local scheduler and a sophisticated remote scheduler to move tasks across worker threads.
Dana Akhmetova, Gokcen Kestor, Roberto Gioiosa, Stefano Markidis, Erwin Laure
CLUSTER5
2015 Evaluation of Parallel Communication Models in Nekbone, a Nek5000 Mini-Application
abstract
Nekbone is a proxy application of Nek5000, a scalable Computational Fluid Dynamics (CFD) code used for modelling incompressible flows. The Nekbone mini-application is used by several international co-design centers to explore new concepts in computer science and to evaluate their performance. We present the design and implementation of a new communication kernel in the Nekbone mini-application with the goal of studying the performance of different parallel communication models. First, a new MPI blocking communication kernel has been developed to solve Nekbone problems in a three-dimensional Cartesian mesh and process topology. The new MPI implementation delivers a 13% performance improvement compared to the original implementation. The new MPI communication kernel consists of approximately 500 lines of code against the original 7,000 lines of code, allowing experimentation with new approaches in Nekbone parallel communication. Second, the MPI blocking communication in the new kernel was changed to the MPI non-blocking communication. Third, we developed a new Partitioned Global Address Space (PGAS) communication kernel, based on the GPI-2 library. This approach reduces the synchronization among neighbor processes and is on average 3% faster than the new MPI-based, non-blocking, approach. In our tests on 8,192 processes, the GPI-2 communication kernel is 3% faster than the new MPI non-blocking communication kernel. In addition, we have used the OpenMP in all the versions of the new communication kernel. Finally, we highlight the future steps for using the new communication kernel in the parent application Nek5000.
Ilya Ivanov, Dana Akhmetova, Ivy Bo Peng, Stefano Markidis, Erwin Laure, Mirko Rahn, Valeria Bartsch, Alistair Hart, Paul F. Fischer
CLUSTER6
2015 The Cost of Synchronizing Imbalanced Processes in Message Passing Systems
abstract
Synchronization in message passing systems is achieved by communication among processes. System and architectural noise and different workloads cause processes to be imbalanced and to reach synchronization points at different time. Thus, both communication and imbalance impact the synchronization performance. In this paper, we study the algorithmic properties that allow the communication in synchronization to absorb the initial imbalance among processes. We quantify the imbalance absorption properties of different barrier algorithms using a LogP Monte Carlo simulator. We found that linear and f-way tournament barriers can absorb up to 95% of random exponential imbalance with the standard deviation equal to the communication time for one message. Dissemination, butterfly and pairwise exchange barriers, on the other hand, do not absorb imbalance but can effectively bound the post-barrier imbalance. We identify that synchronization transits from communication-dominated to imbalance-dominated when the standard deviation of imbalance distribution is more than twice the communication time for one message. In our study, f-way tournament barriers provided the best imbalance absorption rate and convenient communication time.
Ivy Bo Peng, Stefano Markidis, Erwin Laure
CLUSTER3
2015 B2SHARE: An Open eScience Data Sharing Platform
abstract
Scientific data sharing is becoming an essential service for data driven science and can significantly improve the scientific process by making reliable, and trustworthy data available. Thereby reducing redundant work, and providing insights on related research and recent advancements. For data sharing services to be useful in the scientific process, they need to fulfill a number of requirements that cover not only discovery, and access to data. But to ensure the integrity, and reliability of published data as well. B2SHARE, developed by the EUDAT project, provides such a data sharing service to scientific communities. For communities that wish to download, install and maintain their own service, it is also available as software. B2SHARE is developed with a focus on user-friendliness, reliability, and trustworthiness, and can be customized for different organizations and use-cases. In this paper we discuss the design, architecture, and implementation of B2SHARE. We show its usefulness in the scientific process with some case studies in the biodiversity field.
Sarah Berenji Ardestani, Carl Johan Hakansson, Erwin Laure, Ilja Livenson, Pavel Stranák, Emanuel Dima, Dennis Blommesteijn, Mark van de Sanden
e-Science3
2015 Automatic On-Line Detection of MPI Application Structure with Event Flow Graphs
Xavier Aguilar, Karl Fürlinger, Erwin Laure
Euro-Par3
2014 MPI Trace Compression Using Event Flow Graphs
Xavier Aguilar, Karl Fürlinger, Erwin Laure
Euro-Par3
2013 ScaBIA: Scalable Brain Image Analysis in the Cloud
Gert Svensson, Erwin Laure, Matthias Eickhoff, Goetz Brasche
CLOSER3
2013 Using Iterative MapReduce for Parallel Virtual Screening
abstract
Virtual Screening is a technique in chemo informatics used for Drug discovery by searching large libraries of molecule structures. Virtual Screening often uses SVM, a supervised machine learning technique used for regression and classification analysis. Virtual screening using SVM not only involves huge datasets, but it is also compute expensive with a complexity that can grow at least up to O(n2). SVM based applications most commonly use MPI, which becomes complex and impractical with large datasets. As an alternative to MPI, MapReduce, and its different implementations, have been successfully used on commodity clusters for analysis of data for problems with very large datasets. Due to the large libraries of molecule structures in virtual screening, it becomes a good candidate for MapReduce. In this paper we present a MapReduce implementation of SVM based virtual screening, using Spark, an iterative MapReduce programming model. We show that our implementation has a good scaling behaviour and opens up the possibility of using huge public cloud infrastructures efficiently for virtual screening.
Laeeq Ahmed, Åke Edlund, Erwin Laure, Ola Spjuth
CloudCom (2)3
2013 Topic 6: Grid, Cluster and Cloud Computing - (Introduction)
Erwin Laure, Odej Kao, Rosa M. Badia, Laurent Lefèvre, Beniamino Di Martino, Radu Prodan, Matteo Turilli, Daniel Warneke
Euro-Par1
2013 Scalability analysis of Dalton, a molecular structure program
Xavier Aguilar, Michael Schliephake, Olav Vahtras, Judit Giménez, Erwin Laure
Future Gener. Comput. Syst.5
2013 Preface
Erwin Laure, Sverker Holmgren
Future Gener. Comput. Syst.1
2011 Scaling Dalton, A Molecular Electronic Structure Program
abstract
Dalton is a molecular electronic structure program featuring common methods of computational chemistry that are based on pure quantum mechanics (QM) as well as hybrid quantum mechanics/molecular mechanics (QM/MM). It is specialized and has a leading position in calculation of molecular properties with a large world-wide user community (over 2000 licenses issued). In this paper, we present a characterization and performance optimization of Dalton that increases the scalability and parallel efficiency of the application. We also propose a solution that helps to avoid the master/worker design of Dalton to become a performance bottleneck for larger process numbers and increase the parallel efficiency.
Xavier Aguilar, Michael Schliephake, Olav Vahtras, Judit Giménez, Erwin Laure
eScience5
2010 Perspectives on grid computing
Uwe Schwiegelshohn, Rosa M. Badia, Marian Bubak, Marco Danelutto, Schahram Dustdar, Fabrizio Gagliardi, Alfred Geiger, Ladislav Hluchý, Dieter Kranzlmüller, Erwin Laure, Thierry Priol, Alexander Reinefeld, Michael M. Resch, Andreas Reuter 0001, Otto Rienhoff, Thomas Rüter, Peter M. A. Sloot, Domenico Talia, Klaus Ullmann, Ramin Yahyapour
Future Gener. Comput. Syst.10
2009 Using Standards-Based Interfaces to Share Data across Grid Infrastructures
abstract
Data grids, such as the ones used by the high energy physics community, are used to share vast amounts of data across geographic locations. However, interactions with grid data are generally limited by the interfaces provided by the corresponding grid’s infrastructure. The standardization of grid interfaces is one way to expand the reach of grid data seamlessly for users as well as to broaden the set of exploitable grid tools. This in turn enables new collaboration possibilities. The Open Grid Forum has created standards related to accessing grid data. In our work, we explore the usability of two OGF standards, RNS and ByteIO, to enable access to data resources residing in EGEE grids. Grids that implement the RNS specification can address named entities in other grids while grids that implement the ByteIO specification can manipulate data associated with named resources in other grids. Data management functionality in EGEE grids was developed before these specifications were created. As such, implementing RNS and ByteIO for EGEE grids is a test of whether these standards can be applied to an existing grid infrastructure. Through the development of the SNARL and SABLE web services, we demonstrate that such an implementation is possible and measure its performance ramifications. These services expand the access to EGEE grid data by enabling interoperability with other standard compliant grid infrastructures.
Karolina Sarnowska-Upton, Andrew S. Grimshaw, Erwin Laure
ICPP3
2009 Interoperation of world-wide production e-Science infrastructures
abstract
Abstract Many production Grid and e‐Science infrastructures have begun to offer services to end‐users during the past several years with an increasing number of scientific applications that require access to a wide variety of resources and services in multiple Grids. Therefore, the Grid Interoperation Now—Community Group of the Open Grid Forum—organizes and manages interoperation efforts among those production Grid infrastructures to reach the goal of a world‐wide Grid vision on a technical level in the near future. This contribution highlights fundamental approaches of the group and discusses open standards in the context of production e‐Science infrastructures. Copyright © 2009 John Wiley & Sons, Ltd.
Morris Riedel, Erwin Laure, Thomas Soddemann, Laurence Field, John-Paul Navarro, James Casey, Maarten Litmaath, Jean-Philippe Baud, Birger Koblitz, Charles E. Catlett, Dane Skow, Cindy Zheng, Philip M. Papadopoulos, Mason J. Katz, Neha Sharma 0001, Oxana Smirnova, Balázs Kónya, Peter W. Arzberger, Frank Würthwein, Abhishek Singh Rana, Terrence Martin, M. Wan, Von Welch, Tony Rimovsky, Steven J. Newhouse, Andrea Vanni, Yoshio Tanaka, Yusuke Tanimura, Tsutomu Ikegami, David Abramson 0001, Colin Enticott, Graham Jenkins, Ruth Pordes, Steven Timm, Gidon Moont, Mona Aggarwal, Dave Colling, Olivier van der Aa, Alex Sim, Vijaya Natarajan, Arie Shoshani, Junmin Gu, Gerson Galang, Riccardo Zappi, Luca Magnoni, Vincenzo Ciaschini, Michele Pace, Valerio Venturi, Moreno Marzolla, Paolo Andreetto, Robert Cowles, Shaowen Wang 0001, Yuji Saeki, Hitoshi Sato, Satoshi Matsuoka, Putchong Uthayopas, Somsak Sriprayoonsakul, Oscar Koeroo, Matthew Viljoen, Laura Pearlman, Stephen Pickles, David Wallom, Glenn Moloney, Jerome Lauret, Jim Marsteller, Paul Sheldon, Surya Pathak, Shaun De Witt, Jirí Mencák, Jens Jensen, Matt Hodges, Derek Ross, Sugree Phatanapherom, Gilbert Netzer, Anders Rhod Gregersen, Mike Jones 0002, Péter Kacsuk, Achim Streit, Daniel Mallmann, Felix Wolf 0001, Thomas Lippert, Thierry Delaitre, Eduardo Huedo, Neil Geddes
Concurr. Comput. Pract. Exp.2
2009 Grid Deployment Experiences: Grid Interoperation
Laurence Field, Erwin Laure, Markus W. Schulz
J. Grid Comput.2
2008 Preface
Massimo Lamanna, Erwin Laure
J. Grid Comput.2
2006 Data Management in Production Grids - Challenges and Techniques
abstract
Summary form only given. Advances in networking and distributed computing allowed the establishment of production grid infrastructures. Today, large-scale production grid infrastructures such as EGEE in Europe, OSG in the US, and NAREGI in Japan are offering their services to many scientific and industrial applications, from domains as diverse as astronomy, biomedicine, computational chemistry, earth sciences, financial simulations, and high energy physics. Grid infrastructures provide these applications a new means for collaborative research by facilitating the sharing of computational and data resources at an unprecedented scale. The efficient and secure sharing of data resources, which can reach several Tera- to Petabytes in some application domains, is one of the main challenges for grid infrastructures. In this article we discuss the main challenges for data sharing on grid infrastructures, present several techniques that are already established on production grid infrastructures or currently being developed and point out main open research issues. We also present examples of application usage of one of the largest grid infrastructures, EGEE
Erwin Laure
SSDBM1
2005 Performance engineering in data Grids
abstract
Abstract The vision of Grid computing is to facilitate worldwide resource sharing among distributed collaborations. With the help of numerous national and international Grid projects, this vision is becoming reality and Grid systems are attracting an ever increasing user base. However, Grids are still quite complex software systems whose efficient use is a difficult and error‐prone task. In this paper we present performance engineering techniques that aim to facilitate an efficient use of Grid systems, in particular systems that deal with the management of large‐scale data sets in the tera‐ and petabyte range (also referred to as data Grids). These techniques are applicable at different layers of a Grid architecture and we discuss the tools required at each of these layers to implement them. Having discussed important performance engineering techniques, we investigate how major Grid projects deal with performance issues particularly related to data Grids and how they implement the techniques presented. Copyright © 2005 John Wiley & Sons, Ltd.
Erwin Laure, Heinz Stockinger, Kurt Stockinger
Concurr. Pract. Exp.1
2005 File-based replica management
Peter Z. Kunszt, Erwin Laure, Heinz Stockinger, Kurt Stockinger
Future Gener. Comput. Syst.2
2004 Replica Management in the European DataGrid Project
David G. Cameron, James Casey, Leanne Guy, Peter Z. Kunszt, Sophie Lemaitre, Gavin McCance, Heinz Stockinger, Kurt Stockinger, Giuseppe Andronico, William H. Bell, Itzhak Ben-Akiva, Diana Bosio, Radovan Chytracek, Andrea Domenici, Flavia Donno, Wolfgang Hoschek, Erwin Laure, Levi Lucio, A. Paul Millar, Livio Salconi, Ben Segal, Mika Silander
J. Grid Comput.17
2004 Preface
Erwin Laure
J. Grid Comput.1
2001 OpusJava: A Java framework for distributed high performance computing
Erwin Laure
Future Gener. Comput. Syst.1
2000 On the implementation of the Opus coordination language
abstract
Opus is a new programming language designed to assist in coordinating the execution of multiple, independent program modules. With the help of Opus, coarse grained task parallelism between data parallel modules can be expressed in a clean and structured way. In this paper we address the problems of how to build a compilation and runtime support system that can efficiently implement the Opus constructs. Our design considers the often-conflicting goals of efficiency and modular construction through software re-use. In particular, we present the system requirements for an efficient Opus implementation, the Opus runtime system, and describe how they work together to provide the underlying services that the Opus compiler needs for a broad class of machines. Copyright © 2000 John Wiley & Sons, Ltd.
Erwin Laure, Matthew Haines, Piyush Mehrotra, Hans P. Zima
Concurr. Pract. Exp.1
1999 ParBlocks - A New Methodology for Specifying Concurrent Method Executions in Opus
Erwin Laure
Euro-Par1
1999 Compiling Data Parallel Tasks for Coordinated Execution
Erwin Laure, Matthew Haines, Piyush Mehrotra, Hans P. Zima
Euro-Par1