Vasilis Bountris

dblp:338/8423 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-8682-7302ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Low-Level I/O Monitoring for Scientific Workflows
abstract
While detailed resource usage monitoring is possible on the low-level using proper tools, associating such usage with higher-level abstractions in the application layer that actually cause the resource usage in the first place presents a number of challenges. Suppose a large-scale scientific data analysis workflow is run using a distributed execution environment such as a compute cluster or cloud environment and we want to analyze the I/O behaviour of it to find and alleviate potential bottlenecks. Different tasks of the workflow can be assigned to arbitrary compute nodes and may even share the same compute nodes. Thus, locally observed resource usage is not directly associated with the individual workflow tasks. By acquiring resource usage profiles of the involved nodes, we seek to correlate the trace data to the workflow and its individual tasks. To accomplish that, we select the proper set of metadata associated with low-level traces that let us associate them with higher-level task information obtained from log files of the workflow execution as well as the job management using a task orchestrator such as Kubernetes with its container management. Ensuring a proper information chain allows the classification of observed I/O on a logical task level and may reveal the most costly or inefficient tasks of a scientific workflow that are most promising for optimization.
Joel Witzke, Ansgar Lößer, Vasilis Bountris, Tobias Wies, Florian Schintke, Björn Scheuermann 0001
ISPDC3
2026 A systematic evaluation of the potential of carbon-aware execution for scientific workflows
abstract
Scientific workflows are widely used to automate scientific data analysis and often involve computationally intensive processing of large datasets on compute clusters. As such, their execution tends to be long-running and resource-intensive, resulting in substantial energy consumption and, depending on the energy mix, carbon emissions. Meanwhile, a wealth of carbon-aware computing methods have been proposed, yet little work has focused specifically on scientific workflows, even though they present a substantial opportunity for carbon-aware computing because they are often significantly delay tolerant, efficiently interruptible, highly scalable and widely heterogeneous. In this study, we first exemplify the problem of carbon emissions associated with running scientific workflows, and then show the potential for carbon-aware workflow execution. For this, we estimate the carbon footprint of seven real-world Nextflow workflows executed on different cluster infrastructures using both average and marginal carbon intensity data. Furthermore, we systematically evaluate the impact of carbon-aware temporal shifting, and the pausing and resuming of the workflow. Moreover, we apply resource scaling to workflows and workflow tasks. Finally, we report the potential reduction in overall carbon emissions, with temporal shifting capable of decreasing emissions by over 80%, and resource scaling capable of decreasing emissions by 67%.
Kathleen West, Youssef Moawad, Fabian Lehmann, Vasilis Bountris, Ulf Leser, Yehia El-khatib, Lauritz Thamsen
Future Gener. Comput. Syst.4
2025 HyProv: Hybrid Provenance Management for Scientific Workflows
abstract
Provenance plays a crucial role in scientific workflow execution, for instance by providing data for failure analysis, real-time monitoring, or statistics on resource utilization for rightsizing allocations. The workflows themselves, however, become increasingly complex in terms of involved components. Furthermore, they are executed on distributed cluster infrastructures, which makes the real-time collection, integration, and analysis of provenance data challenging. Existing provenance systems struggle to balance scalability, real-time processing, online provenance analytics, and integration across different components and compute resources. Moreover, most provenance solutions are not workflow-aware; by focusing on arbitrary workloads, they miss opportunities for workflow systems where optimization and analysis can exploit the availability of a workflow specification that dictates, to some degree, task execution orders and provides abstractions for physical tasks at a logical level. In this paper, we present HyProv, a hybrid provenance management system that combines centralized and federated paradigms to offer scalable, online, and workflow-aware queries over workflow provenance traces. HyProv uses a centralized component for efficient management of the small and stable workflow-specification-specific provenance, and complements this with federated querying over different scalable monitoring and provenance databases for the large-scale execution logs. This enables low-latency access to current execution data. Furthermore, the design supports complex provenance queries, which we exemplify for the workflow system Airflow in combination with the resource manager Kubernetes. Our experiments indicate that HyProv scales to large workflows, answers provenance queries with sub-second latencies, and adds only modest CPU and memory overhead to the cluster.
Vasilis Bountris, Lauritz Thamsen, Ulf Leser
IEEE Big Data1
2025 Domain-Specific Data Compression for Nextflow with COMET-FLOW
Ninon De Mecquenem, Simon Bosse, Vasilis Bountris, Fabian Lehmann, Somayeh Mohammadi, Pauline Karega, Knut Reinert, Ulf Leser
IEEE Big Data3
2025 Exploring the Potential of Carbon-Aware Execution for Scientific Workflows
abstract
Scientific workflows are widely used to automate scientific data analysis and often involve processing large quantities of data on compute clusters. As such, their execution tends to be long-running and resource intensive, leading to significant energy consumption and carbon emissions. Meanwhile, a wealth of carbon-aware computing methods have been proposed, yet little work has focused specifically on scientific workflows, even though they present a substantial opportunity for carbon-aware computing because they are inherently delay tolerant, efficiently interruptible, and highly scalable. In this study, we demonstrate the potential for carbonaware workflow execution. For this, we estimate the carbon footprint of two real-world Nextflow workflows executed on cluster infrastructure. We use a linear power model for energy consumption estimates and real-world average and marginal CI data for two regions. We evaluate the impact of carbonaware temporal shifting, pausing and resuming, and resource scaling. Our findings highlight significant potential for reducing emissions of workflows and workflow tasks.
Kathleen West, Fabian Lehmann, Vasilis Bountris, Ulf Leser, Yehia El-khatib, Lauritz Thamsen
CCGrid3
2024 Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
abstract
Scientific workflows are used to analyze large amounts of data. These workflows comprise numerous tasks, many of which are executed repeatedly, running the same custom program on different inputs. Users specify resource allocations for each task, which must be sufficient for all inputs to prevent task failures. As a result, task memory allocations tend to be overly conservative, wasting precious cluster resources, limiting overall parallelism, and increasing workflow makespan.In this paper, we first benchmark a state-of-the-art method on four real-life workflows from the nf-core workflow repository. This analysis reveals that certain assumptions underlying current prediction methods, which typically were evaluated only on simulated workflows, cannot generally be confirmed for real workflows and executions. We then present Ponder, a new online task-sizing strategy that considers and chooses between different methods to cater to different memory demand patterns. We implemented Ponder for Nextflow and made the code publicly available. In an experimental evaluation that also considers the impact of memory predictions on scheduling, Ponder improves Memory Allocation Quality on average by 71.0% and makespan by 21.8% in comparison to a state-of-the-art method. Moreover, Ponder produces 93.8% fewer task failures.
Fabian Lehmann, Jonathan Bader, Ninon De Mecquenem, Vasilis Bountris, Florian Friederici, Ulf Leser, Lauritz Thamsen
e-Science5
2024 CuttleFlow: Infrastructure-Specific Workflow Adaption for Improved Reusability
abstract
Scientific workflows have gained popularity for large-scale data analysis due to their potential to improve the reproducibility, scalability and documentation of complex multistep scientific analysis pipelines. However, their reusability is currently limited in practice, as a workflow is typically developed for a specific infrastructure. This is reflected in the choice of tools (e.g. less/more memory requirements), their configuration (e.g. number of threads) and the workflow topology (e.g. data parallel scatter/gather). Re-running such a workflow requires access to the same, or at least a highly similar, computing environment, effectively reducing its use by other groups. To address this challenge, we present CuttleFlow, a novel method for adapting and rewriting scientific workflows given a description of an infrastructure and its inputs. CuttleFlow starts from an abstract workflow description and compiles it into an infrastructure-specific logical workflow using three types of rewriting operations, namely tool replacement, tool reconfiguration, and data scattering/gathering for task parallelization. We implement a prototype based on NextFlow and evaluate it for two important bioinformatics data analysis problems, namely RNAseq and metagenomics, on a distributed infrastructure. We demonstrate the large impact that the rewriting of CuttleFlow can have on runtime, achieving a reduction in makespan of up to 71%. We also demonstrate a significant reduction in resource usage through our rewriting approach.
Ninon De Mecquenem, Simon Bosse, Vasilis Bountris, Somayeh Mohammadi, Knut Reinert, Ulf Leser
e-Science3
2024 I/O of Scientific Workflows Monitored in Detail
abstract
Correlating detailed local resource utilization data with the high-level concepts of distributed scientific workflow systems eventually causing it is challenging. When running a large-scale scientific data analysis workflow across a distributed execution environment, we want to analyze its I/O behaviour to identify potential bottlenecks. Since tasks are assigned to any available nodes, local resource usage on a node does not directly show which tasks are causing it. We acquire resource usage profiles of the involved nodes to link them to the individual workflow tasks. This is done by properly associating low-level trace metadata with high-level task information from log files and job management systems like Kubernetes. This information helps identifying areas of the workflow on a logical task level where improvements can make the biggest impact.
Joel Witzke, Ansgar Lößer, Vasilis Bountris, Florian Schintke, Björn Scheuermann 0001
e-Science3
2022 Are Workflows a Language to Solve Software Management Challenges? - A ßMACH Based Analysis
abstract
A language in computer science is not just a tool, the language defines essential aspects of the software management process. The language determines persons/actors/roles in the development/management team, the type of solution strategy, the description of the project, the cooperation with other teams, and the definition of the process. We are using workflows: The workflow language defines a problem decomposition to sub-workflows and tasks. We analyzed this language by a systematic method to nail down the software management process definition and to describe all relevant aspects of software management. This method is ßMACH and based on ßMACH we can provide the finding that the workflow language strongly defines a wide set of management aspects. As a result, we can state that the impact of a language is non-negligible.
Marcus Hilbrich, Vasilis Bountris
SoMeT2