Spiros Koulouzis

dblp:03/8770 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
6since 2021 · last 2023
0000-0001-8652-315XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Systems, architecture and hardware · 7 · 3 first-authorComputer networks · 1
YearPublicationVenuePosition
2023 Towards a Service-based Adaptable Data Layer for Cloud Workflows
abstract
Many scientific workflows are data-driven and need to be continuously executed for the large volume of datasets transferred from distributed data sources. The overhead arising from data transfers must be considered when optimizing workflow performance. Many workflow systems support various data transfer protocols (DTPs) and file systems. However, challenges that hinder wide protocol adoption are mainly the need for more feasibility of adapting new solutions, such as decentralized ones. In this paper, we prototype a container-native data layer that supports multiple DTPs, e.g., FTP, WebDAV, and IPFS, for Cloud workflows. Based on this tool, we demonstrated the feasibility of using combinations of Docker, CWL, and Argo to deploy and execute several application scenarios adaptably. Besides, we analyzed the performance of data transfers and workflow execution time between IPFS and WebDAV, which can help users decide which one to handle data. Our results show that IPFS outperforms WebDAV in uploading large files, and the makespan via IPFS executed in Argo is comparable with WebDAV.
Yuandou Wang, Nikita Janse, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
COMPSAC4
2023 Integrating R in a Distributed Scientific Workflow via a Jupyter-Based Environment
abstract
The Research Infrastructure Lifewatch Italy has developed a Virtual Research Environment for studies on phytoplankton ecology that includes computational services based on R, a programming language widely used for data science and ecology. Here we have verified the feasibility of a Jupyter-based research environment, the NaaVRE, which has so far been tested only with Python, for running R code in a workflow on the Cloud. The successful execution demonstrated the potentialities of R in Cloud-based research environments. However, further investigation is needed, in particular, to overcome the issue of the lack of dependencies declaration in R. The possibility of performing analyses in a workflow, combined with the computational resources of remote infrastructures, will support scientists in carrying out FAIR and innovative research in a more efficient, integrated and collaborative way.
Mariantonietta La Marra, Diederik Blanson Henkemans, Jessica Titocci, Spiros Koulouzis, Ilaria Rosati, Zhiming Zhao
e-Science4
2022 Context-Aware Notebook Search in a Jupyter-Based Virtual Research Environment
abstract
Computational notebook environments such as the Jupyter play an increasingly important role in data-centric research for prototyping computational experiments, documenting code implementations, and sharing scientific results. Effectively discovering and reusing notebooks available on the web can reduce repetitive work and facilitate scientific innovations. However, general-purpose web search engines (e.g., Google Search) do not explicitly index the contents of notebooks, and notebook repositories (e.g., Kaggle and GitHub) require users to create domain-specific queries based on the metadata in the notebook catalogs, which fail to capture the working contexts in the notebook environment. This poster presents a Context-aware Notebook Search Framework (CANSF) to enable a researcher to seamlessly discover external notebooks based on semantic contexts of the literate programming activities in the Jupyter environment.
Siamak Farshidi, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
e-Science4
2022 Featured Cover
abstract
The cover image is based on the Research Article Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment by Zhiming Zhao et al., https://doi.org/10.1002/spe.3098.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.2
2022 Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment
abstract
Abstract Virtual research environments (VREs) provide user‐centric support in the lifecycle of research activities, for example, discovering and accessing research assets or composing and executing application workflows. A typical VRE is often implemented as an integrated environment, including a catalog of research assets, a workflow management system, a data management framework, and tools for enabling user collaboration. In contrast, notebook environments like Jupyter allow researchers to rapidly prototype scientific code and share their experiments as online accessible notebooks. Jupyter can support several popular languages used by data scientists, such as Python, R, and Julia. However, such notebook environments do not have seamless support for running heavy computations on remote infrastructure or finding and accessing collaborative software code inside notebooks. This article investigates the gap between a notebook environment and a VRE and proposes an embedded VRE solution for the Jupyter environment called Notebook‐as‐a‐VRE (NaaVRE). The NaaVRE solution provides functional components via a component marketplace and allows users to create a customized VRE on top of the Jupyter environment. From the VRE, a user can search research assets (data, software, and algorithms), compose workflows, manage the lifecycle of an experiment, and share the results among users in the community. We demonstrate how such a solution can enhance a legacy workflow that uses Light Detection and Ranging (LiDAR) data from country‐wide airborne laser scanning surveys for deriving geospatial data products of ecosystem structure at high resolution over broad spatial extents. This enables users to scale out the processing of multi‐terabyte LiDAR point clouds for ecological applications to more data sources in a distributed cloud environment. Similar applications could be developed for workflows producing other essential biodiversity variables.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.2
2021 The ARTICONF approach to decentralized car-sharing
abstract
Social media applications are essential for next-generation connectivity. Today, social media are centralized platforms with a single proprietary organization controlling the network and posing critical trust and governance issues over the created and propagated content. The ARTICONF project funded by the European Union's Horizon 2020 program researches a decentralized social media platform based on a novel set of trustworthy, resilient and globally sustainable tools that address privacy, robustness and autonomy-related promises that proprietary social media platforms have failed to deliver so far. This paper presents the ARTICONF approach to a car-sharing decentralized application (DApp) use case, as a new collaborative peer-to-peer model providing an alternative solution to private car ownership. We describe a prototype implementation of the car-sharing social media DApp and illustrate through real snapshots how the different ARTICONF tools support it in a simulated scenario.
Nishant Saurabh, Carlos Rubia, Anandakumar Palanisamy, Spiros Koulouzis, Mirsat Sefidanoski, Antorweep Chakravorty, Zhiming Zhao, Aleksandar Karadimce, Radu Prodan
Blockchain Res. Appl.4
2020 Decentralized Social Media Applications as a Service: a Car-Sharing Perspective
abstract
Social media applications are essential for next generation connectivity. Today, social media are centralized platforms with a single proprietary organization controlling the network and posing critical trust and governance issues over the created and propagated content. The ARTICONF project funded by the European Union’s Horizon 2020 program researches a decentralized social media platform based on a novel set of trustworthy, resilient and globally sustainable tools to fulfil the privacy, robustness and autonomy-related promises that proprietary social media platforms have failed to deliver so far. This paper presents the ARTICONF approach to a car-sharing use case application, as a new collaborative peer-to-peer model providing an alternative solution to private car ownership. We describe a prototype implementation of the car-sharing social media application and illustrate through real snapshots how the different ARTICONF tools support it in a simulated scenario.
Anandhakumar Palanisamy, Mirsat Sefidanoski, Spiros Koulouzis, Carlos Rubia, Nishant Saurabh, Radu Prodan
ISCC3
2020 Time-critical data management in clouds: Challenges and a Dynamic Real-Time Infrastructure Planner (DRIP) solution
abstract
Summary The increasing volume of data being produced, curated, and made available by research infrastructures in the environmental science domain require services that are able to optimize the delivery and staging of data for researchers and other users of scientific data. Specialized data services for managing data life cycle, for creating and delivering data products, and for customized data processing and analysis all play a crucial role in how these research infrastructures serve their communities, and many of these activities are time‐critical—needing to be carried out frequently within specific time windows. We describe our experiences identifying the time‐critical requirements of environmental scientists making use of computational research support environments. We present a microservice‐based infrastructure optimization suite, the Dynamic Real‐Time Infrastructure Planner, used for constructing virtual infrastructures for research applications on demand. We provide a case study whereby our suite is used to optimize runtime service quality for a data subscription service provided by the Euro‐Argo using EGI Federated Cloud and EUDAT's B2SAFE services, and to consider how such a case study relates to other application scenarios.
Spiros Koulouzis, Paul Martin 0002, Huan Zhou 0006, Yang Hu 0013, Thierry Carval, Baptiste Grenier, Jani Heikkinen, Cees T. A. M. de Laat, Zhiming Zhao
Concurr. Comput. Pract. Exp.1
2019 Contextual Linking between Workflow Provenance and System Performance Logs
abstract
When executing scientific workflows, anomalies of the workflow behavior are often caused by different issues such as resource failures at the underlying infrastructure. The provenance information collected by workflow management systems only captures the transformation of data at the workflow level. Analyzing provenance information and apposite system metrics requires expertise and manual effort. Moreover, it is often time-consuming to aggregate this information and correlate events occurring at different levels of the infrastructure. In this paper, we propose an architecture to automate the integration among workflow provenance information and performance information from the infrastructure level. Our architecture enables workflow developers or domain scientists to effectively browse workflow execution information together with the system metrics, and analyze contextual information for possible anomalies.
Elias el Khaldi Ahanach, Spiros Koulouzis, Zhiming Zhao
eScience2
2019 Teaching DevOps and Cloud Based Software Engineering in University Curricula
abstract
This paper presents recommendations on the design and pilot implementation of the DevOps and Cloud based Software Development curricula for Computer Science and Software Engineering masters. The central part of proposed approach is the Body of Knowledge in the DevOps technologies for Software Engineering (DevOpsSE BoK) that defines a set Knowledge Areas and Knowledge Units required for SE professionals to work efficiently as DevOps engineer or application developer. Defining DevOpsSE-BoK provides a basis for defining required professional competences and skills and allows consistent curricula structuring and profiling. The paper also reports on the experience of the first course run on 2018/2019 academic year at the University of Amsterdam. The paper presents the structure of the course and explains what instructional methodologies have been used for course development, such as project based learning that facilitates the students' team based skills both in mastering Agile development process and skills sharing. The paper provides a short summary of the generally used DevOps definitions, concepts, models and tools, specifically focusing on the cloud based DevOps tools for software development, deployment and operation that allows the main DevOps principle of continuous development and continuous improvement which are critical for modern agile data driven companies.
Yuri Demchenko, Zhiming Zhao, Jayachander Surbiryala, Spiros Koulouzis, Zeshun Shi, Jelena Gordiyenko
eScience4
2019 Bridging the demand and the offer in data science
abstract
Summary During the last several years, we have observed an exponential increase in the demand for Data Scientists in the job market. As a result, a number of trainings, courses, books, and university educational programs (both at undergraduate, graduate and postgraduate levels) have been labeled as “Big data” or “Data Science”; the fil‐rouge of each of them is the aim at forming people with the right competencies and skills to satisfy the business sector needs. In this paper, we report on some of the exercises done in analyzing current Data Science education offer and matching with the needs of the job markets to propose a scalable matching service, ie, COmpetencies ClassificatiOn (E‐CO‐2), based on Data Science techniques. The E‐CO‐2 service can help to extract relevant information from Data Science–related documents (course descriptions, job Ads, blogs, or papers), which enable the comparison of the demand and offer in the field of Data Science Education and HR management, ultimately helping to establish the profession of Data Scientist.
Adam Belloum, Spiros Koulouzis, Tomasz Wiktorski, Andrea Manieri
Concurr. Comput. Pract. Exp.2
2019 SWITCH workbench: A novel approach for the development and deployment of time-critical microservice-based cloud-native applications
Polona Stefanic, Matej Cigale, Andrew C. Jones, Louise Knight, Ian J. Taylor, Cristiana Istrate, George Suciu, Alexandre Ulisses, Vlado Stankovski, Salman Taherizadeh, Guadalupe Flores Salado, Spiros Koulouzis, Paul Martin 0002, Zhiming Zhao
Future Gener. Comput. Syst.12
2019 CloudsStorm: A framework for seamlessly programming and controlling virtual infrastructure functions during the DevOps lifecycle of cloud applications
abstract
Summary The infrastructure‐as‐a‐service (IaaS) model of cloud computing provides virtual infrastructure functions (VIFs), which allow application developers to flexibly provision suitable virtual machines' (VM) types and locations, and even configure the network connection for each VM. Because of the pay‐as‐you‐go business model, IaaS provides an elastic way to operate applications on demand. However, in current cloud applications DevOps (software development and operations) lifecycle, the VM provisioning steps mainly rely on manually leveraging these VIFs. Moreover, these functions cannot be programmatically embedded into the application logic to control the infrastructure at runtime. Especially, the vendor lock‐in issue, which different clouds provide different VIFs, also enlarges this gap between the cloud infrastructure management and application operation. To mitigate this gap, we designed and implemented a framework, CloudsStorm, which enables developers to easily leverage VIFs of different clouds and program them into their cloud applications. To be specific, CloudsStorm empowers applications with infrastructure programmability at design‐level, infrastructure‐level, and application‐level. CloudsStorm also provides two infrastructure controlling modes, ie, active and passive mode, for applications at runtime. Besides, case studies about operating task‐based and big data applications on clouds show that the monetary cost is significantly reduced through the seamless and on‐demand infrastructure management provided by CloudsStorm. Finally, the scaling and recovery operation evaluations of CloudsStorm are performed to show its controlling performance. Compared with other tools, ie, “jcloud” and “cloudinit.d”, the scaling and provisioning performance evaluations demonstrate that CloudsStorm can achieve at least 10% efficiency improvement in our experiment settings.
Huan Zhou 0006, Yang Hu 0013, Xue Ouyang 0003, Jinshu Su, Spiros Koulouzis, Cees T. A. M. de Laat, Zhiming Zhao
Softw. Pract. Exp.5
2018 Information Centric Networking for Sharing and Accessing Digital Objects with Persistent Identifiers on Data Infrastructures
abstract
Persistent identifiers (PIDs) such as Digital Object Identifiers (DOIs) provide a unique and persistent way to identify and cite digital objects such as publications, media content and research data. They are widely used by data producers to catalogue and publish digital assets and research data. Nowadays, research infrastructures (RIs) offer services not only for accessing and publishing data objects, but also for processing data based on user demands, e.g., via scientific workflows or third party virtual research environments. However, efficiently retrieving and sharing digital objects in a shared data processing environment requires knowledge of application access patterns as well as the underlying network level distribution. As the number and size of data objects increases, optimizing data discovery and access among distributed partners on shared infrastructure emerges as an important challenge for infrastructure operators to maintain quality of service and user experience. In this paper, we propose a novel approach that utilizes Information Centric Networking (ICN) to retrieve content based on PIDs while optimizing data access on shared infrastructure.
Spiros Koulouzis, Rahaf Mousa, Andreas Karakannas, Cees T. A. M. de Laat, Zhiming Zhao
CCGrid1
2017 Automatic Collector for Dynamic Cloud Performance Information
abstract
When deploying an application in the cloud, a developer often wants to know which of the wide variety of cloud resources is best to use. Most cloud providers only provide static information about different cloud resources which is often not enough because static information does not take into account the hardware and software that is being used or the policy that has been applied by the cloud provider. Therefore, dynamic benchmarking of cloud resources is needed to find out how a certain workload is going to behave on a certain instance. However, benchmarking various cloud resources is a time consuming process. Thus, using a tool which automatically benchmarks various cloud resources will be of great use. In this paper, we present the Cloud Performance Collector, a modular cloud benchmarking tool aimed to automatically benchmark a wide variety of applications. To demonstrate the benefit of the tool, we did three experiments with three synthetic benchmark applications and one real-world application using the ExoGENI testbed.
Olaf Elzinga, Spiros Koulouzis, Arie Taal, Yang Hu 0013, Huan Zhou 0006, Paul Martin 0002, Cees T. A. M. de Laat, Zhiming Zhao
NAS2
2016 SDN-aware federation of distributed data
Spiros Koulouzis, Adam Belloum, Marian Bubak, Zhiming Zhao, Miroslav Zivkovic, Cees T. A. M. de Laat
Future Gener. Comput. Syst.1
2014 Applying workflow as a service paradigm to application farming
abstract
SUMMARY Task farming is often used to enable parameter sweep for exploration of large sets of initial conditions for large scale complex simulations. Such applications occur very often in life sciences. Available solutions enable to perform parameter sweep by creating multiple job submissions with different parameters. This paper presents an approach to farm workflows, employing service oriented paradigms using the WS‐VLAM workflow manager, which provides ways to create, control, and monitor workflow applications and their components. We present two service‐oriented approaches for workflow farming: task level, whereby task harness acts as services by being invoked on which task to load, and data level, where the actual task is invoked as a service with different chunks of data to process. An experimental evaluation of the presented solution is performed with a biomedical application for which 3000 simulations were required to perform a Monte Carlo study. Copyright © 2013 John Wiley & Sons, Ltd.
Reginald Cushing, Spiros Koulouzis, Adam Belloum, Marian Bubak
Concurr. Comput. Pract. Exp.2
2011 Dynamic Handling for Cooperating Scientific Web Services
abstract
Many e-Science applications are increasingly relying on orchestrating workflows of static web services. The static nature of these web services means that workflow management systems have no control over the underlying mechanics of such services. This lack of control manifests itself as a problem when optimizing workflow execution since techniques such as data-locality aware deployment and service-to-service communication are very difficult to achieve. In this paper we propose a novel approach for mobilizing scientific web services onto common distributed resources and as such enable back-to-back communication between cooperating web services, autonomous web service scaling through fuzzy control and autonomous web service workflow orchestration.
Reginald Cushing, Spiros Koulouzis, Adam Belloum, Marian Bubak
eScience2
2008 Enabling Data Transport between Web Services through alternative protocols and Streaming
abstract
A significant challenge for Web services is to find reliable and efficient methods to transfer large data between them. This paper describes the problem of scalable data transport between Web services, and proposes a solution: the development of a modular server/client library that uses SOAP as a control channel while the actual data transport is accomplished by various protocol implementation. Apart from file transport, the proposed approach offers the facility of direct data streaming between Web services, an approach that could benefit workflow execution time.
Spiros Koulouzis, Edgar Meij, M. Scott Marshall, Adam Belloum
eScience1