Ninon De Mecquenem

dblp:331/5389 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-3052-6129ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Domain-Specific Data Compression for Nextflow with COMET-FLOW
Ninon De Mecquenem, Simon Bosse, Vasilis Bountris, Fabian Lehmann, Somayeh Mohammadi, Pauline Karega, Knut Reinert, Ulf Leser
IEEE Big Data1
2025 WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
abstract
Scientific workflows process extensive data sets over clusters of independent nodes, which requires a complex stack of infrastructure components, especially a resource manager (RM) for task-to-node assignment, a distributed file system (DFS) for data exchange between tasks, and a workflow engine to control task dependencies. To enable a decoupled development and installation of these components, current architectures place intermediate data files during workflow execution independently of the future workload. In data-intensive applications, this separation results in suboptimal schedules, as tasks are often assigned to nodes lacking input data, causing network traffic and bottlenecks. This paper presents WOW, a new scheduling approach for dynamic scientific workflow systems that steers both data movement and task scheduling to reduce network congestion and overall runtime. For this, WOW creates speculative copies of intermediate files to prepare the execution of subsequently scheduled tasks. WOW supports modern workflow systems that gain flexibility through the dynamic construction of execution plans. We prototypically implemented WOW for the popular workflow engine Nextflow using Kubernetes as a resource manager. In experiments with 16 synthetic and real workflows, WOW reduced makespan in all cases, with improvement of up to 94.5 % for workflow patterns and up to 53.2 % for real workflows, at a moderate increase of temporary storage space. It also has favorable effects on CPU allocation and scales well with increasing cluster size.
Fabian Lehmann, Jonathan Bader, Friedrich Tschirpke, Ninon De Mecquenem, Ansgar Lößer, Sören Becker 0001, Katarzyna Ewa Lewinska, Lauritz Thamsen, Ulf Leser
CCGrid4
2024 Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
abstract
Scientific workflows are used to analyze large amounts of data. These workflows comprise numerous tasks, many of which are executed repeatedly, running the same custom program on different inputs. Users specify resource allocations for each task, which must be sufficient for all inputs to prevent task failures. As a result, task memory allocations tend to be overly conservative, wasting precious cluster resources, limiting overall parallelism, and increasing workflow makespan.In this paper, we first benchmark a state-of-the-art method on four real-life workflows from the nf-core workflow repository. This analysis reveals that certain assumptions underlying current prediction methods, which typically were evaluated only on simulated workflows, cannot generally be confirmed for real workflows and executions. We then present Ponder, a new online task-sizing strategy that considers and chooses between different methods to cater to different memory demand patterns. We implemented Ponder for Nextflow and made the code publicly available. In an experimental evaluation that also considers the impact of memory predictions on scheduling, Ponder improves Memory Allocation Quality on average by 71.0% and makespan by 21.8% in comparison to a state-of-the-art method. Moreover, Ponder produces 93.8% fewer task failures.
Fabian Lehmann, Jonathan Bader, Ninon De Mecquenem, Vasilis Bountris, Florian Friederici, Ulf Leser, Lauritz Thamsen
e-Science3
2024 CuttleFlow: Infrastructure-Specific Workflow Adaption for Improved Reusability
abstract
Scientific workflows have gained popularity for large-scale data analysis due to their potential to improve the reproducibility, scalability and documentation of complex multistep scientific analysis pipelines. However, their reusability is currently limited in practice, as a workflow is typically developed for a specific infrastructure. This is reflected in the choice of tools (e.g. less/more memory requirements), their configuration (e.g. number of threads) and the workflow topology (e.g. data parallel scatter/gather). Re-running such a workflow requires access to the same, or at least a highly similar, computing environment, effectively reducing its use by other groups. To address this challenge, we present CuttleFlow, a novel method for adapting and rewriting scientific workflows given a description of an infrastructure and its inputs. CuttleFlow starts from an abstract workflow description and compiles it into an infrastructure-specific logical workflow using three types of rewriting operations, namely tool replacement, tool reconfiguration, and data scattering/gathering for task parallelization. We implement a prototype based on NextFlow and evaluate it for two important bioinformatics data analysis problems, namely RNAseq and metagenomics, on a distributed infrastructure. We demonstrate the large impact that the rewriting of CuttleFlow can have on runtime, achieving a reduction in makespan of up to 71%. We also demonstrate a significant reduction in resource usage through our rewriting approach.
Ninon De Mecquenem, Simon Bosse, Vasilis Bountris, Somayeh Mohammadi, Knut Reinert, Ulf Leser
e-Science1
2024 Validity constraints for data analysis workflows
abstract
Porting a scientific data analysis workflow (DAW) to a cluster infrastructure, a new software stack, or even only a new dataset with some notably different properties is often challenging. Despite the structured definition of the steps (tasks) and their interdependencies during a complex data analysis in the DAW specification, relevant assumptions may remain unspecified and implicit. Such hidden assumptions often lead to crashing tasks without a reasonable error message, poor performance in general, non-terminating executions, or silent wrong results of the DAW, to name only a few possible consequences. Searching for the causes of such errors and drawbacks in a distributed compute cluster managed by a complex infrastructure stack, where DAWs for large datasets typically are executed, can be tedious and time-consuming. We propose validity constraints (VCs) as a new concept for DAW languages to alleviate this situation. A VC is a constraint specifying logical conditions that must be fulfilled at certain times for DAW executions to be valid. When defined together with a DAW, VCs help to improve the portability, adaptability, and reusability of DAWs by making implicit assumptions explicit. Once specified, VCs can be controlled automatically by the DAW infrastructure, and violations can lead to meaningful error messages and graceful behaviour (e.g., termination or invocation of repair mechanisms). We provide a broad list of possible VCs, classify them along multiple dimensions, and compare them to similar concepts one can find in related fields. We also provide a proof-of-concept implementation for the workflow system Nextflow.
Florian Schintke, Khalid Belhajjame, Ninon De Mecquenem, David Frantz, Vanessa Emanuela Guarino, Marcus Hilbrich, Fabian Lehmann, Paolo Missier, Rebecca Sattler, Jan Arne Sparka, Daniel T. Speckhard, Hermann Stolte, Duc Anh Vu 0001, Ulf Leser
Future Gener. Comput. Syst.3
2023 Design by Contract Revisited in the Context of Scientific Data Analysis Workflows
abstract
Software systems enabling large-scale data analysis workflows (DAWs) are a key technology for modern science as they allow extracting new insights from experimental results. DAWs are pipelines composed of interdependent tasks that are executed in a distributed fashion on large compute clusters. Typically, the individual task implementations are developed by research groups all over the world and usually not tested outside a narrow scope of possible inputs, parameters, and infrastructures. As a result, the operations' correctness depends on many implicit assumptions, such as the completeness and suitability of input data, infrastructure properties such as available cores, etc. This makes quality assurance of DAWs a critical issue. We propose to address this problem by introducing a contract-driven approach to DAW design and implementation. Following the well-known principle of Design by Contract, DAW developers specify contracts in the form of requirements and promises for each task of a DAW. These contracts serve as guards to ensure that tasks run in a proper environment and produce correct results. The detection of contract violations allows to halt the execution of a DAW and to identify the culprit that caused the violation. Thus, the integration of contracts into DAW design provides opportunities for efficiency improvements by reducing computation and debugging time in case of errors.
Duc Anh Vu 0001, Jan Arne Sparka, Ninon De Mecquenem, Timo Kehrer, Ulf Leser, Lars Grunske
e-Science3
2023 Contract-Driven Design of Scientific Data Analysis Workflows
abstract
Software systems enabling large-scale data analysis workflows (DAWs) are a key technology for many scientific disciplines, as they allow extracting new insights from experimental results. DAWs are (non-)linear pipelines composed of multiple interdependent tasks that are executed in a distributed fashion on large compute clusters. In science, the individual task implementations are developed by research groups all over the world and usually not tested outside a narrow scope of possible inputs, parameters, and infrastructures. As a result, the operations' correctness depends on many implicit assumptions. Among others this includes the completeness and suitability of input data, infrastructure properties such as available cores or main memory, etc. This combination of complexity, distribution and untested components makes quality assurance of DAWs a critical issue. In this paper, we propose to address this problem by introducing a contract-driven approach to DAW design and implementation. Following this method, DAW developers specify contracts in the form of requirements and promises for each task of a DAW. These contracts serve as guards to ensure that tasks run in a proper environment and produce correct results. We provide the first formal definition of contracts for DAWs and show how they are connected to DAW scheduling and execution. As a proof of concept, we extended Nextflow, a popular scientific workflow system, with contracts and defined a light-weight DSL for their specification. We exemplify the power of a contract-driven approach to DAW development by enhancing several real-world DAWs from Bioinformatics to capture typical problems during their execution and show how the specific notifications issued by broken contracts help debugging the DAWs.
Duc Anh Vu 0001, Jan Arne Sparka, Ninon De Mecquenem, Timo Kehrer, Ulf Leser, Lars Grunske
e-Science3
2023 A mathematical programming approach for resource allocation of data analysis workflows on heterogeneous clusters
abstract
Abstract Scientific communities are motivated to schedule their large-scale data analysis workflows in heterogeneous cluster environments because of privacy and financial issues. In such environments containing considerably diverse resources, efficient resource allocation approaches are essential for reaching high performance. Accordingly, this research addresses the scheduling problem of workflows with bag-of-task form to minimize total runtime (makespan). To this aim, we develop a mixed-integer linear programming model (MILP). The proposed model contains binary decision variables determining which tasks should be assigned to which nodes. Also, it contains linear constraints to fulfill the tasks requirements such as memory and scheduling policy. Comparative results show that our approach outperforms related approaches in most cases. As part of the post-optimality analysis, some secondary preferences are imposed on the proposed model to obtain the most preferred optimal solution. We analyze the relaxation of the makespan in the hope of significantly reducing the number of consumed nodes.
Somayeh Mohammadi, Latif Pourkarimi, Felix Droop, Ninon De Mecquenem, Ulf Leser, Knut Reinert
J. Supercomput.4
2022 A Consolidated View on Specification Languages for Data Analysis Workflows
Marcus Hilbrich, Sebastian Müller 0007, Svetlana Kulagina, Christopher Lazik, Ninon De Mecquenem, Lars Grunske
ISoLA (2)5