EDBT 2026 Demo / reviewers in the wild / expert
Frédéric Suter
dblp:66/3341 · also Fred Suter
· DBLP profile ↗
46ranked-venue papers
7as first author
18since 2021 · last 2027
0000-0003-1902-1955ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DTLMod: A simulation framework for in situ workflow optimization
Frédéric Suter |
Future Gener. Comput. Syst. | 1 |
| 2026 | JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific WorkflowsabstractIn modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. Vladislav Esaulov, Jieyang Chen, Norbert Podhorszki, Frédéric Suter, Scott Klasky, Anu G. Bourgeois, Lipeng Wan 0001 |
HPDC | 4 |
| 2026 | Alternative mixed integer linear programming optimization for joint job scheduling and data allocation in grid computingabstractThis paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as a mixed integer quadratically constrained program. To tackle the nonlinearity in the constraint, we alternatively fix a subset of decision variables and optimize the remaining ones via Mixed Integer Linear Programming (MILP). We solve the MILP problem at each iteration via an off-the-shelf MILP solver. Our experimental results show that our method significantly outperforms existing heuristic methods, employing either independent optimization or joint optimization strategies. We have also verified the generalization ability of our method over grid environments with various sizes and its high robustness to the algorithm setting. Shengyu Feng, Jaehyung Kim 0001, Yiming Yang 0002, Joseph Boudreau, Tasnuva Chowdhury, Adolfy Hoisie, Raees Khan, Ozgur O. Kilic, Scott Klasky, Tatiana Korchuganova, Paul Nilsson, Verena Ingrid Martinez Outschoorn, David Keetae Park, Norbert Podhorszki, Yihui Ren 0001, Frédéric Suter, Sairam Sri Vatsavai, Shinjae Yoo, Tadashi Maeno, Alexei Klimentov |
Future Gener. Comput. Syst. | 16 |
| 2026 | A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva |
Future Gener. Comput. Syst. | 1 |
| 2025 | A Versatile Simulated Data Transport Layer for in Situ Workflows Performance EvaluationabstractIn situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes. Frédéric Suter |
CLUSTER | 1 |
| 2025 | A Performance Model of In-Situ Techniques
Yi Ju, Nicolas Vidal 0003, Adalberto Perez, Ana Gainaru, Frédéric Suter, Stefano Markidis, Philipp Schlatter, Scott Klasky, Erwin Laure |
PDP | 5 |
| 2025 | Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid
Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, Frédéric Suter |
Parallel Comput. | 5 |
| 2024 | Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AIabstractAs Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions. Renan Souza 0001, Silvina Caíno-Lores, Mark Coletti, Tyler J. Skluzacek, Alexandru Costan, Frédéric Suter, Marta Mattoso, Rafael Ferreira da Silva |
e-Science | 6 |
| 2024 | Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy ScienceabstractNuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows. Frédéric Suter, Norbert Podhorszki, Scott Klasky |
e-Science | 1 |
| 2024 | Automated Calibration of a Simulator of MPI Application ExecutionsabstractThe traditional approach for assessing the performance of scientific applications on HPC platforms consists in executing these applications on these platforms. But conducting these real-world experiments comes with several difficulties. Besides being often time-, labor-, and resource-intensive, experiments are limited to application and platform configurations at hand, thus precluding the exploration of "what if?" scenarios. A way to resolve these difficulties is to resort to simulation. The main concern, then, is that of simulation accuracy. For a simulation to be accurate, the parameters that define the behavior of the simulation models can be calibrated with respect to ground-truth executions. Simulation calibration, in the current state of the art, relies, at best, on labor-intensive manual procedures. We propose an automated simulation calibration approach, and apply this approach to the specific context of the simulation of MPI applications on leadership class HPC platforms. This poster will motivate the development of this approach and detail our methodology and results. Yick Ching Wong, Frédéric Suter, Kshitij Mehta, Henri Casanova, Jesse McDonald |
e-Science | 2 |
| 2023 | Accelerating the Design of Scientific Workflows with Simulation-Based Rapid PrototypingabstractExtreme-scale science requires scientists to combine multiple heterogeneous computational tasks in complex workflows, efficiently managing large amounts of data, and fully exploiting the performance of the entire edge-to-HPC computing continuum. Going from science to workflows requires domain scientists to express their research ideas from a computer science perspective so that the most adapted and efficient tools and techniques can be selected. However, this is a long, error-prone, and tedious process, as there is no “turnkey solution” in such a diverse ecosystem. Therefore, we propose to develop a comprehensive simulation-based framework that will be used by domain scientists to easily prototype their scientific workflows, while expressing all the important information needed by computer scientists to provide them with the most efficient implementation. Frédéric Suter |
e-Science | 1 |
| 2023 | Driving Next-Generation Workflows from the Data PlaneabstractWe observe the emergence of a new generation of scientific workflows that process data produced at a sustained rate by scientific instruments and large scale numerical simulations. This data is consumed by multiple analysis, visualization, or Machine Learning components not only to enable inference and justify the scientific program, but also to monitor and steer the evolution of these experiments. In such workflows, moving intermediate data efficiently is key to performance, more than efficiently scheduling computational tasks. However, most traditional workflow management systems focus on optimizing task scheduling and then deal with data management, assuming a “move little, compute for long” model, which makes them unfit to the efficient management of this new generation of workflows. Therefore, we advocate for a new way to manage scientific workflows. We propose to consider an efficiently and independently managed data plane that can store and stream data. Workflows compute components, in the application plane can then interact with the data plane, abstracted from complexities of data management. Then, the role of a workflow management system would become that of a control plane that allows users to connect services together to execute the workflow and manages connections between the application and data planes. In this position paper, we characterize several next-generation workflow motifs and describe how their interaction with the data plane is a challenge to traditional workflow management systems. Then, we express a set of requirements that a workflow management system should meet to efficiently manage next-generation workflows at different scales. Based on these requirements, we expose our vision of driving next-generation workflows from the data plane and list remaining open challenges. Frédéric Suter, Rafael Ferreira da Silva, Ana Gainaru, Scott Klasky |
e-Science | 1 |
| 2022 | Hybrid Analysis of Fusion Data for Online Understanding of Complex Science on Extreme Scale ComputersabstractThe current practice for fusion scientists running first principle simulations on high performance computing plat-forms is to either run their simulations and output their data for post-hoc analysis, or to place in situ analytics into their code. In this paper we examine a complex workflow using XGC fusions simulation run on the Oak Ridge Leadership Computing Facility's supercomputer Summit, which also involve three anal-yses as part of the results necessary for scientific discovery. We discuss the challenges faced when implementing these algorithms and present an original hybrid staging technique to help enable the physicists to make discoveries during the execution of the simulation. By creating this infrastructure, we can examine complicated physics results, which may not have been possible without the infrastructure. For example, our work enables the online visualization of turbulent homoclinic tangle around the magnetic X-point, breaking the last confinement surface. This visualization could help fusion scientists to better understand and improve the turbulence spread of plasma exhaust heat, which is crucial toward realizing plasmas beyond the currently accessible physics regimes of present-day tokamak reactors. The physics of turbulent homoclinic tangle will be reported in a future physics publication, by utilizing the original online analysis/visualization framework presented in this paper. Eric Suchyta, Jong Choi 0001, Seung-Hoe Ku, David Pugmire, Ana Gainaru, Kevin A. Huck, Ralph Kube, Aaron Scheinberg, Frédéric Suter, Choong-Seock Chang, Todd S. Munson, Norbert Podhorszki, Scott Klasky |
CLUSTER | 9 |
| 2022 | SIM-SITU: A Framework for the Faithful Simulation of in situ ProcessingabstractThe amount of data generated by numerical simulations in various scientific domains led to a fundamental redesign of how the analysis and visualization of simulation outputs are performed. The throughput and capacity of storage subsystems have not evolved as fast as the computing power in extreme-scale supercomputers, making the classical post-hoc approach highly inefficient. In situ processing has then emerged as a solution in which simulation and data analysis/visualization are intertwined for better performance and greater interactivity. Determining the best allocation, i.e., how many resources to allocate to simulation and analysis respectively, mapping, i.e., where and at which frequency to run the analysis/visualization, and data transfer mode is a complex task whose performance assessment is crucial to the efficient execution of in situ processing. However, such a performance evaluation of different strategies usually relies either on directly running them on the targeted execution environments, which can rapidly become extremely time- and resource-consuming, or on resorting to simplified models of the components of an in situ application, which can lack of realism. In both cases, the validity of the performance evaluation is limited. In this paper, we present SIM-SITU, a framework for the faithful performance evaluation of in situ processing strategies. We designed SIM-SITU to reflect the typical features of in situ processing systems. Thanks to its modular design, Sim-Situ has the necessary flexibility to easily and faithfully evaluate the behavior and performance of various allocation, mapping, and data transfer strategies. We illustrate the capabilities of SIM-SITU on a Molecular Dynamics use case. We study the impact of different strategies on performance and show how users can leverage SIM-SITU to determine interesting tradeoffs when adding analysis/visualization components to their application. Valentin Honoré, Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Ewa Deelman, Frédéric Suter |
e-Science | 6 |
| 2022 | Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path ForwardabstractThe ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows. Kshitij Mehta, Ashley Cliff, Frédéric Suter, Angelica M. Walker, Matthew Wolf, Daniel A. Jacobson, Scott Klasky |
e-Science | 3 |
| 2022 | Data-aware and simulation-driven planning of scientific workflows on IaaS cloudsabstractAbstract The promise of an easy access to a virtually unlimited number of resources makes Infrastructure as a Service Clouds a good candidate for the execution of data‐intensive workflow applications composed of hundreds of computational tasks. Thanks to a careful execution planning, workflow management systems can build a tailored compute infrastructure by combining a set of virtual machine instances. However, these applications usually rely on files to handle dependencies between tasks. A storage space shared by all virtual machines may become a bottleneck and badly impact the application execution time. In this article, we propose an original data‐aware planning algorithm that leverages two characteristics of a family of virtual machines instances, that is, a large number of cores and a dedicated storage space on fast SSD drives, to improve data locality, hence reducing the amount of data transfers over the network during the execution of a workflow. We also propose a simulation‐driven approach to solve a cost‐performance optimization problem and correctly dimension the virtual infrastructure onto which execute a given workflow. Experiments conducted with real application workflows show the benefits of the presented algorithms. The data‐aware planning leads to a clear reduction of both execution time and volume of data transferred over the network while the simulation‐driven approach allows us to dimension the infrastructure in a reasonable time. Tchimou N'Takpé, Jean Edgard Gnimassoun, Souleymane Oumtanaga, Frédéric Suter |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Sequence-RTG: Efficient and Production-Ready Pattern Mining in System Log MessagesabstractSystem logs are a wealth of information that can be leveraged to control the behaviour of a computing and storage infrastructure, detect deviations from normal behaviour, and react accordingly by triggering some predefined actions. System log management usually consists of a complex workflow that collects, standardises, indexes, stores, and visualises the log messages to help system administration teams in their daily operations. In large scale data centres such log management infrastructures can collect millions if not billions of messages per day. A key component in this workflow is the identification of message patterns, which requests the expertise of administrators. These patterns represent a template of both static and variable message parts against which a new log message can be matched. This crucial task is often done manually, but these patterns can change frequently making it time consuming for the human operators to keep up. Therefore, we propose in this paper to automate the discovery of patterns in system log messages by extending the functionalities of an existing pattern mining framework, called Sequence. Our main objectives are to improve both the scalability of this framework and its capacity to be integrated into a complete system log management workflow. We present how we addressed six main limitations of the seminal Sequence tool. These modifications led us to propose Sequence-RTG (Sequence-Ready-To-Go), a more efficient and production-ready version. We analyse its performance in terms of both speed, using data-sets of increasing sizes, and accuracy on data-sets from the literature. We also show that two months after the introduction of Sequence-RTG within the system log management framework of the IN2P3 Computing Centre we reduced the fraction of messages that are not matched to a pattern from 75-80% to only 15%. Louise Harding, Fabien Wernli, Frédéric Suter |
CLUSTER | 3 |
| 2021 | Learning-Based Approaches to Estimate Job Wait Time in HTC Datacenters
Luc Gombert, Frédéric Suter |
JSSPP | 2 |
| 2020 | Developing accurate and scalable simulators of production workflow management systems with WRENCH
Henri Casanova, Rafael Ferreira da Silva, Ryan Tanaka, Suraj Pandey, Gautam Jethwani, William Koch, Spencer Albrecht, James Oeth, Frédéric Suter |
Future Gener. Comput. Syst. | 9 |
| 2019 | Bridging Concepts and Practice in eScience via Simulation-Driven EngineeringabstractThe CyberInfrastructure (CI) has been the object of intensive research and development in the last decade, resulting in a rich set of abstractions and interoperable software implementations that are used in production today for supporting ongoing and breakthrough scientific discoveries. A key challenge is the development of tools and application execution frameworks that are robust in current and emerging CI configurations, and that can anticipate the needs of upcoming CI applications. This paper presents WRENCH, a framework that enables simulation-driven engineering for evaluating and developing CI application execution frameworks. WRENCH provides a set of high-level simulation abstractions that serve as building blocks for developing custom simulators. These abstractions rely on the scalable and accurate simulation models that are provided by the SimGrid simulation framework. Consequently, WRENCH makes it possible to build, with minimum software development effort, simulators that that can accurately and scalably simulate a wide spectrum of large and complex CI scenarios. These simulators can then be used to evaluate and/or compare alternate platform, system, and algorithm designs, so as to drive the development of CI solutions for current and emerging applications. Rafael Ferreira da Silva, Henri Casanova, Ryan Tanaka, Frédéric Suter |
eScience | 4 |
| 2019 | Improving Fairness in a Large Scale HTC System Through Workload Analysis and Simulation
Frédéric Azevedo, Dalibor Klusácek, Frédéric Suter |
Euro-Par | 3 |
| 2018 | Reducing the Human-in-the-Loop Component of the Scheduling of Large HTC Workloads
Frédéric Azevedo, Luc Gombert, Frédéric Suter |
JSSPP | 3 |
| 2017 | Modeling Distributed Platforms from Application Traces for Realistic File Transfer SimulationabstractSimulation is a fast, controlled, and reproducible way to evaluate new algorithms for distributed computing platforms in a variety of conditions. However, the realism of simulations is rarely assessed, which critically questions the applicability of a whole range of findings. In this paper, we present our efforts to build platform models from application traces, to allow for the accurate simulation of file transfers across a distributed infrastructure. File transfers are key to performance, as the variability of file transfer times has important consequences on the dataflow of the application. We present a methodology to build realistic platform models from application traces and provide a quantitative evaluation of the accuracy of the derived simulations. Results show that the proposed models are able to correctly capture real-life variability and significantly outperform the state-of-the-art model. Anchen Chai, Mohammad-Mahdi Bazm, Sorina Camarasu-Pop, Tristan Glatard, Hugues Benoit-Cattin, Frédéric Suter |
CCGrid | 6 |
| 2017 | Don't Hurry Be Happy: A Deadline-Based Backfilling Approach
Tchimou N'Takpé, Frédéric Suter |
JSSPP | 2 |
| 2017 | Simulating MPI Applications: The SMPI ApproachabstractThis article summarizes our recent work and developments on SMPI, a flexible simulator of MPI applications. In this tool, we took a particular care to ensure our simulator could be used to produce fast and accurate predictions in a wide variety of situations. Although we did build SMPI on SimGrid whose speed and accuracy had already been assessed in other contexts, moving such techniques to a HPC workload required significant additional effort. Obviously, an accurate modeling of communications and network topology was one of the key to such achievements. Another less obvious key was the choice to combine in a single tool the possibility to do both offline and online simulation. Augustin Degomme, Arnaud Legrand, George S. Markomanolis, Martin Quinson, Mark Stillwell, Frédéric Suter |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2015 | Adding Storage Simulation Capacities to the SimGrid Toolkit: Concepts, Models, and APIabstractFor each kind of distributed computing infrastructures, i.e., clusters, grids, clouds, data centers, or supercomputers, storage is a essential component to cope with the tremendous increase in scientific data production and the ever-growing need for data analysis and preservation. Understanding the performance of a storage subsystem or dimensioning it properly is an important concern for which simulation can help by allowing for fast, fully repeatable, and configurable experiments for arbitrary hypothetical scenarios. However, most simulation frameworks tailored for the study of distributed systems offer no or little abstractions or models of storage resources. In this paper, we detail the extension of SimGrid, a versatile toolkit for the simulation of large-scale distributed computing systems, with storage simulation capacities. We first define the required abstractions and propose anew API to handle storage components and their contents in SimGrid-based simulators. Then we characterize the performance of the fundamental storage component that are disks and derive models of these resources. Finally we list several concrete use cases of storage simulations in clusters, grids, clouds, and data centers for which the proposed extension would be beneficial. Adrien Lèbre, Arnaud Legrand, Frédéric Suter, Pierre Veyre |
CCGRID | 3 |
| 2015 | Simulation of MPI applications with time-independent tracesabstractSummary Analyzing and understanding the performance behavior of parallel applications on parallel computing platforms is a long‐standing concern in the High Performance Computing community. When the targeted platforms are not available, simulation is a reasonable approach to obtain objective performance indicators and explore various hypothetical scenarios. In the context of applications implemented with the Message Passing Interface, two simulation methods have been proposed, on‐line simulation and off‐line simulation, both with their own drawbacks and advantages. In this work, we present an off‐line simulation framework, that is, one that simulates the execution of an application based on event traces obtained from an actual execution. The main novelty of this work, when compared to previously proposed off‐line simulators, is that traces that drive the simulation can be acquired on large, distributed, heterogeneous, and non‐dedicated platforms. As a result, the scalability of trace acquisition is increased, which is achieved by enforcing that traces contain no time‐related information. Moreover, our framework is based on a state‐of‐the‐art scalable, fast, and validated simulation kernel. We introduce the notion of performing off‐line simulation from time‐independent traces, propose and evaluate several trace acquisition strategies, describe our simulation framework, and assess its quality in terms of trace acquisition scalability, simulation accuracy, and simulation time. Copyright © 2014 John Wiley & Sons, Ltd. Henri Casanova, Frédéric Desprez, George S. Markomanolis, Frédéric Suter |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | Versatile, scalable, and accurate simulation of distributed applications and platforms
Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, Frédéric Suter |
J. Parallel Distributed Comput. | 5 |
| 2012 | Scalable Multi-purpose Network Representation for Large Scale Distributed System SimulationabstractConducting experiments in large-scale distributed systems is usually time-consuming and labor-intensive. Uncontrolled external load variation prevents to reproduce experiments and such systems are often not available to the purpose of research experiments, e.g. production or yet to deploy systems. Hence, many researchers in the area of distributed computing rely on simulation to perform their studies. However, the simulation of large-scale computing systems raises several scalability issues, in terms of speed and memory. Indeed, such systems now comprise millions of hosts interconnected through a complex network and run billions of processes. Most simulators thus trade accuracy for speed and rely on very simple and easy to implement models. However, the assumptions underlying these models are often questionable, especially when it comes to network modeling. In this paper, we show that, despite a widespread belief in the community, achieving high scalability does not necessarily require to resort to overly simple models and ignore important phenomena. We show that relying on a modular and hierarchical platform representation, while taking advantage of regularity when possible, allows us to model systems such as data and computing centers, peer-to-peer networks, grids, or clouds in a scalable way. This approach has been integrated into the open-source SimGrid simulation toolkit. We show that our solution allows us to model such systems much more accurately than other state-of-the-art simulators without trading for simulation speed. SimGrid is even sometimes orders of magnitude faster. Laurent Bobelin, Arnaud Legrand, David A. González Márquez, Pierre Navarro, Martin Quinson, Frédéric Suter, Christophe Thiery |
CCGRID | 6 |
| 2012 | Budget Constrained Resource Allocation for Non-deterministic Workflows on an IaaS Cloud
Eddy Caron, Frédéric Desprez, Adrian Muresan, Frédéric Suter |
ICA3PP (1) | 4 |
| 2011 | Introduction
Leonel Sousa, Frédéric Suter, Alfredo Goldman, Rizos Sakellariou, Oliver Sinnen |
Euro-Par (1) | 2 |
| 2011 | Single Node On-Line Simulation of MPI Applications with SMPIabstractSimulation is a popular approach for predicting the performance of MPI applications for platforms that are not at one's disposal. It is also a way to teach the principles of parallel programming and high-performance computing to students without access to a parallel computer. In this work we present SMPI, a simulator for MPI applications that uses on-line simulation, i.e., the application is executed but part of the execution takes place within a simulation component. SMPI simulations account for network contention in a fast and scalable manner. SMPI also implements an original and validated piece-wise linear model for data transfer times between cluster nodes. Finally SMPI simulations of large-scale applications on large-scale platforms can be executed on a single node thanks to techniques to reduce the simulation's compute time and memory footprint. These contributions are validated via a large set of experiments in which SMPI is compared to popular MPI implementations with a view to assess its accuracy, scalability, and speed. Pierre-Nicolas Clauss, Mark Stillwell, Stéphane Genaud, Frédéric Suter, Henri Casanova, Martin Quinson |
IPDPS | 4 |
| 2010 | A Bi-criteria Algorithm for Scheduling Parallel Task Graphs on ClustersabstractApplications structured as parallel task graphs exhibit both data and task parallelism, and arise in many domains. Scheduling these applications on parallel platforms has been a long-standing challenge. In the case of a single homogeneous cluster, most of the existing algorithms focus on the reduction of the application completion time (make span). But in presence of resource managers such as batch schedulers and due to accentuated pressure on energy concerns, the produced schedules also have to be efficient in terms of resource usage. In this paper we propose a novel bi-criteria algorithm, called biCPA, able to optimize these two performance metrics either simultaneously or separately. Using simulation over a wide range of experimental scenarios, we find that biCPA leads to better results than previously published algorithms. Frédéric Desprez, Frédéric Suter |
CCGRID | 2 |
| 2010 | Minimizing Stretch and Makespan of Multiple Parallel Task Graphs via Malleable AllocationsabstractMany scientific applications can be structured as Parallel Task Graphs (PTGs), i.e., graphs of data-parallel tasks. Adding data-parallelism to a task-parallel application provides opportunities for higher performance and scalability, but poses scheduling challenges. We study the off-line scheduling of multiple PTGs on a single, homogeneous cluster. The objective is to optimize performance and fairness. We propose a novel algorithm that first computes perfectly fair PTG completion times assuming that each PTG is an ideal malleable job. These completion times are then relaxed so that the schedule is organized as a sequence of periods and is still close to the perfectly fair schedule. Finally, since PTGs are not perfectly malleable, the algorithm increases the execution time of all PTGs uniformly until it can successfully schedule each task in a period. Our evaluation in simulation, using both synthetic and real-world application configurations, shows that our algorithm outperforms previously proposed algorithms when considering two different performance metrics and one fairness metric. Henri Casanova, Frédéric Desprez, Frédéric Suter |
ICPP | 3 |
| 2010 | On cluster resource allocation for multiple parallel task graphs
Henri Casanova, Frédéric Desprez, Frédéric Suter |
J. Parallel Distributed Comput. | 3 |
| 2009 | Concurrent scheduling of parallel task graphs on multi-clusters using constrained resource allocationsabstractScheduling multiple applications on heterogeneous multi-clusters is challenging as the different applications have to compete for resources. A scheduler thus has to ensure a fair distribution of resources among the applications and prevent harmful selfish behaviors while still trying to minimize their respective completion time. In this paper we consider mixed-parallel applications, represented by graphs whose nodes are data-parallel tasks, that are scheduled in two steps: allocation and mapping. We investigate several strategies to constrain the amount of resources the scheduler can allocate to each application and evaluate them over a wide range of scenarios. Tchimou N'Takpé, Frédéric Suter |
IPDPS | 2 |
| 2009 | Scheduling Parallel Task Graphs on (Almost) Homogeneous Multicluster PlatformsabstractApplications structured as parallel task graphs exhibit both data and task parallelism and arise in many domains. Scheduling these applications efficiently on parallel platforms has been a long-standing challenge. In the case of a single homogeneous platform, such as a cluster, results have been obtained both in theory, i.e., guaranteed algorithms, and, in practice, i.e., pragmatic heuristics. Due to task parallelism, these applications are well suited for execution on distributed platforms that span multiple clusters possibly in multiple institutions. However, the only available results in this context are nonguaranteed heuristics. In this paper, we develop a scheduling algorithm, MCGAS, which is applicable to multicluster platforms that are almost homogeneous. Such platforms are often found as large subsets of multicluster platforms. Our novel contribution is that MCGAS computes task allocations so that a (tunable) performance guarantee is provided. Since a performance guarantee does not necessarily imply good average performance in practice, we also compare MCGAS with a recently proposed nonguaranteed algorithm. Using simulation over a wide range of experimental scenarios, we find that MCGAS leads to better average application makespans than its competitor. Pierre-François Dutot, Tchimou N'Takpé, Frédéric Suter, Henri Casanova |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | Scheduling Dynamic Workflows onto Clusters of Clusters using PostponingabstractIn this article, we revisit the problem of scheduling dynamically generated directed acyclic graphs (DAGs) of multi-processor tasks (M-tasks). A DAG is a basic model for expressing workflows applications where each node represents a task of the workflow. We present a novel algorithm (DMHEFT) for scheduling dynamically generated DAGs onto a heterogeneous collection of clusters. The scheduling decisions are based on the predicted runtime of an M-task as well as the estimation of the redistribution costs between data-dependent tasks. The algorithm also takes care of unfavorable placements of M-tasks by considering the postponing of ready tasks even if idle processors are available. We evaluate the scheduling algorithm by comparing the resulting makespans to the results obtained by using other scheduling algorithms, such as RePA and MHEFT. Sascha Hunold, Thomas Rauber, Frédéric Suter |
CCGRID | 3 |
| 2008 | Redistribution aware two-step scheduling for mixed-parallel applicationsabstractApplications raising in many scientific fields exhibit both data and task parallelism that have to be exploited efficiently. A classic approach is to structure those applications by a task graph whose nodes represent parallel computations. Scheduling such mixed-parallel applications is challenging even on a single homogeneous platform, such as a cluster. Most of the mixed-parallel application scheduling algorithms rely on two decoupled steps: allocation and mapping. This separation can induce unnecessary or costly data redistributions that have an impact on the overall performance. This is particularly true for data intensive applications. In this paper, we propose an original approach in which the allocations determined in the first step can be adapted during the second step in order to minimize the impact of these data redistributions. Two redistribution aware mapping strategies are detailed and a study of their impact on the schedule length is proposed through a comparison with an efficient two step algorithm over a broad range of experimental scenarios. Sascha Hunold, Thomas Rauber, Frédéric Suter |
CLUSTER | 3 |
| 2008 | Out-of-Core Wavefront Computations with Reduced SynchronizationabstractMatrix computation algorithms often exhibit dependencies between neighboring elements inside loop nests such that the frontier between computed elements and those to be computed wanders inform of a 'wave' through the matrix. Macro-pipelining techniques can achieve an efficient parallelization of such algorithms by overlapping communication and computation. Usually these techniques are limited to situations where all the data to be processed fits into main memory, whereas for larger data the I/O usage pattern for external storage requires special attention. The work presented a first extension of the wavefront framework to these so-called out-of-core problems. The present paper proposes a redesign of their algorithm that minimizes both overhead and perturbations coming from communications. To tackle the issue of non-contiguous I/O, we also propose an optimized data layout. These two major modifications of the original algorithm eventually allow us to present a third improvement as our implementation shortens the transition phase between two consecutive iterations of the wavefront algorithm. Experiments performed with the PARXXL library show that we can significantly reduce the time lost during inefficient I/O operations and thus obtain faster computations. Pierre-Nicolas Clauss, Jens Gustedt, Frédéric Suter |
PDP | 3 |
| 2007 | A Comparison of Scheduling Approaches for Mixed-Parallel Applications on Heterogeneous PlatformsabstractMixed-parallel applications can take advantage of large-scale computing platforms but scheduling them efficiently on such platforms is challenging. In this paper we compare the two main proposed approaches for solving this scheduling problem on a heterogeneous set of homogeneous clusters. We first modify previously proposed algorithms for both approaches and show that our modifications lead to significant improvements. We then perform a comparison of the modified algorithms in simulation over a wide range of application and platform conditions. We find that although both approaches have advantages, one of them is most likely the most appropriate for the majority of users. Tchimou N'Takpé, Frédéric Suter, Henri Casanova |
ISPDC | 2 |
| 2004 | From Heterogeneous Task Scheduling to Heterogeneous Mixed Parallel Scheduling
Frédéric Suter, Frédéric Desprez, Henri Casanova |
Euro-Par | 1 |
| 2004 | Impact of mixed-parallelism on parallel implementations of the Strassen and Winograd matrix multiplication algorithmsabstractAbstract In this paper we study the impact of the simultaneous exploitation of data‐ and task‐parallelism, so called mixed‐parallelism, on the Strassen and Winograd matrix multiplication algorithms. This work takes place in the context of Grid computing and, in particular, in the Client–Agent(s)–Server(s) model, where data can already be distributed on the platform. For each of those algorithms, we propose two mixed‐parallel implementations. The former follows the phases of the original algorithms while the latter has been designed as the result of a list scheduling algorithm. We give a theoretical comparison, in terms of memory usage and execution time, between our algorithms and classical data‐parallel implementations. This analysis is corroborated by experiments. Finally, we give some hints about heterogeneous and recursive versions of our algorithms. Copyright © 2004 John Wiley & Sons, Ltd. Frédéric Desprez, Frédéric Suter |
Concurr. Pract. Exp. | 2 |
| 2002 | A Scalable Approach to Network Enabled Servers (Research Note)
Eddy Caron, Frédéric Desprez, Frédéric Lombard, Jean-Marc Nicod, Laurent Philippe 0001, Martin Quinson, Frédéric Suter |
Euro-Par | 7 |
| 2001 | Mixed Parallel Implementations of Strassen and Winograd Matrix Multiplication AlgorithmsabstractThis paper presents parallel implementations of the top level of Strassen and Winograd algorithms for matrix multiplication that use mixed-parallelism, i.e., simultaneous exploitation of data- and task-parallelism. This paradigm allows a better task placement and reduces the communication costs. A comparison with the ScaLAPACK implementation of the matrix multiplication is given. We present a theoretical evaluation of the algorithms which is corroborated by experiments. Frédéric Desprez, Frédéric Suter |
IPDPS | 2 |
| 2001 | SCILAB to SCILAB//: The OURAGAN project
Eddy Caron, Serge Chaumette, Sylvain Contassot-Vivier, Frédéric Desprez, Eric Fleury, Claude Gomez, Maurice Goursat, Martin Quinson, Emmanuel Jeannot, Dominique Lazure, Frédéric Lombard, Jean-Marc Nicod, Laurent Philippe 0001, Pierre Ramet, Jean Roman, Frank Rubi, Serge Steer, Frédéric Suter, Gil Utard |
Parallel Comput. | 18 |