Yuandou Wang

dblp:190/5790 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-4694-9572ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 D-VRE: From a Jupyter-enabled private research environment to decentralized collaborative research ecosystem
abstract
Today, scientific research is increasingly becoming data-centric and compute-intensive, relying on data and models across distributed sources. However, challenges still exist in the traditional cooperation mode, given the high storage and computing costs, geolocation barriers, and local confidentiality regulations. The Jupyter environment has recently emerged and evolved into a vital virtual research environment for scientific computing, which researchers can use to scale computational analyses up to larger datasets and high-performance computing resources. Nevertheless, existing approaches lack robust support of a decentralized cooperation mode to unlock the full potential of decentralized collaborative scientific research, e.g., seamlessly secure data sharing. In this work, we change the basic structure and legacy norms of current research environments via the seamless integration of Jupyter with Ethereum blockchain capabilities. As such, it creates a Decentralized Virtual Research Environment (D-VRE) from private computational notebooks to a decentralized collaborative research ecosystem. We propose a novel architecture for the D-VRE and prototype some essential D-VRE elements for enabling secure data sharing with decentralized identity, user-centric agreement-making, membership, and research asset management. To validate our method, we conduct an experimental study to test all functionalities of D-VRE smart contracts and their gas consumption. In addition, we deploy the D-VRE prototype on a test net of the Ethereum blockchain for demonstration. The feedback from the studies showcases the current prototype's usability, ease of use, and potential, and suggests further improvements.
Yuandou Wang, Sheejan Tripathi, Siamak Farshidi, Zhiming Zhao
Blockchain Res. Appl.1
2024 CrowdAL: Towards a Blockchain-empowered Active Learning System in Crowd Data Labeling
abstract
Active Learning (AL) is a machine learning technique where the model selectively queries the most informative data points for labeling by human experts. Integrating AL with crowdsourcing leverages crowd diversity to enhance data labeling but introduces challenges in consensus and privacy. This poster presents CrowdAL, a blockchain-empowered crowd AL system designed to address these challenges. CrowdAL integrates blockchain for transparency and a tamper-proof incentive mechanism, using smart contracts to evaluate crowd workers’ performance and aggregate labeling results, and employs zeroknowledge proofs to protect worker privacy.
Shaojie Hou, Yuandou Wang, Zhiming Zhao
e-Science2
2024 A Collaborative Framework for Facilitating Federated Learning among Jupyter Users
abstract
Federated learning (FL) allows multiple partners to train machine learning models without sharing raw data, thus preserving privacy. Despite its promising aspects, existing FL frameworks have some drawbacks regarding flexibility, decentralized aggregation, and collaborative environments. This poster presents FedLearn, a collaborative community framework built atop the JupyterLab environment for FL among Jupyter users. We use a microservices architecture to implement the framework and enable automated FL deployment processes across multiple clouds. The demonstration showcases the feasibility of the FedLearn community portal for Jupyter users.
Anandan Krishnasamy, Yuandou Wang, Zhiming Zhao
e-Science2
2024 PriCE: Privacy-Preserving and Cost-Effective Scheduling for Parallelizing the Large Medical Image Processing Workflow over Hybrid Clouds
Yuandou Wang, Neel Kanwal, Kjersti Engan, Chunming Rong, Paola Grosso, Zhiming Zhao
Euro-Par (1)1
2023 Towards a Service-based Adaptable Data Layer for Cloud Workflows
abstract
Many scientific workflows are data-driven and need to be continuously executed for the large volume of datasets transferred from distributed data sources. The overhead arising from data transfers must be considered when optimizing workflow performance. Many workflow systems support various data transfer protocols (DTPs) and file systems. However, challenges that hinder wide protocol adoption are mainly the need for more feasibility of adapting new solutions, such as decentralized ones. In this paper, we prototype a container-native data layer that supports multiple DTPs, e.g., FTP, WebDAV, and IPFS, for Cloud workflows. Based on this tool, we demonstrated the feasibility of using combinations of Docker, CWL, and Argo to deploy and execute several application scenarios adaptably. Besides, we analyzed the performance of data transfers and workflow execution time between IPFS and WebDAV, which can help users decide which one to handle data. Our results show that IPFS outperforms WebDAV in uploading large files, and the makespan via IPFS executed in Argo is comparable with WebDAV.
Yuandou Wang, Nikita Janse, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
COMPSAC1
2023 Towards a Knowledge Graph Enhanced Automation and Collaboration Framework for Digital Twins
abstract
The Digital Twin (DT) provides a digital representation of a physical system and allows users to interactively study the physical processes of a real system via the digital representation in different scenarios in real time. The development of a DT is highly complex; it requires not only expertise from multiple disciplines but also the integration of often heterogeneous software components, e.g., simulations, machine learning, visualization, and user interface components across distributed environments. This poster presents a Knowledge Graph-based ontological framework to boost automation and collaboration during the DT lifecycle stages. We implement our methods in developing a what-if analysis service for a DT of an ecosystem of wetlands and its automated deployment to the Amazon Web Services (AWS) cloud.
Vasileios Christou, Yuandou Wang, Zhiming Zhao
e-Science2
2023 CWL-FLOps: A Novel Method for Federated Learning Operations at Scale
abstract
Federated Learning (FL) has attracted much attention in recent years because it enables users with private data sets to train a global model collaboratively without raw data exchange. However, due to a lack of automation, researchers often struggled to develop, deploy, track, and manage all the data, steps, and configuration setup for all FL participating nodes. Federated Learning Operations (FLOps) is recently emerging in the FL community, a new methodology for developing FL systems efficiently and continuously. Some research works discussed approaches for FLOps, but only a few solutions address managing FL application scenarios from the workflow perspective. This poster proposes CWL-FLOps, a novel CWL-based method for FLOps, which can improve the flexibility of FL abstraction and fully automate the FL deployment and execution by mapping high-level descriptions onto distributed resource nodes. Our experiments demonstrate the feasibility of describing centralized and decentralized FL scenarios using CWL abstracted definitions without relying on heavily customized or external software for execution.
Chronis Kontomaris, Yuandou Wang, Zhiming Zhao
e-Science2
2023 A Survey on Dataset Distillation: Approaches, Applications and Future Directions
abstract
Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, including support for continual learning, neural architecture search, and privacy protection. Despite recent advances, we lack a holistic understanding of the approaches and applications. Our survey aims to bridge this gap by first proposing a taxonomy of dataset distillation, characterizing existing approaches, and then systematically reviewing the data modalities, and related applications. In addition, we summarize the challenges and discuss future directions for this field of research.
Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong
IJCAI3
2022 Featured Cover
abstract
The cover image is based on the Research Article Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment by Zhiming Zhao et al., https://doi.org/10.1002/spe.3098.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.7
2022 Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment
abstract
Abstract Virtual research environments (VREs) provide user‐centric support in the lifecycle of research activities, for example, discovering and accessing research assets or composing and executing application workflows. A typical VRE is often implemented as an integrated environment, including a catalog of research assets, a workflow management system, a data management framework, and tools for enabling user collaboration. In contrast, notebook environments like Jupyter allow researchers to rapidly prototype scientific code and share their experiments as online accessible notebooks. Jupyter can support several popular languages used by data scientists, such as Python, R, and Julia. However, such notebook environments do not have seamless support for running heavy computations on remote infrastructure or finding and accessing collaborative software code inside notebooks. This article investigates the gap between a notebook environment and a VRE and proposes an embedded VRE solution for the Jupyter environment called Notebook‐as‐a‐VRE (NaaVRE). The NaaVRE solution provides functional components via a component marketplace and allows users to create a customized VRE on top of the Jupyter environment. From the VRE, a user can search research assets (data, software, and algorithms), compose workflows, manage the lifecycle of an experiment, and share the results among users in the community. We demonstrate how such a solution can enhance a legacy workflow that uses Light Detection and Ranging (LiDAR) data from country‐wide airborne laser scanning surveys for deriving geospatial data products of ecosystem structure at high resolution over broad spatial extents. This enables users to scale out the processing of multi‐terabyte LiDAR point clouds for ecological applications to more data sources in a distributed cloud environment. Similar applications could be developed for workflows producing other essential biodiversity variables.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.7
2021 Reliability-Aware and Deadline-Constrained Mobile Service Composition Over Opportunistic Networks
abstract
An opportunistic link between two mobile devices or nodes can be constructed when they are within each other’s communication range. Typically, cyber–physical environments consist of a number of mobile devices that are potentially able to establish opportunistic contacts and serve mobile applications in a cost-effective way. Opportunistic mobile service computing is a promising paradigm capable of utilizing the pervasive mobile computational resources around the users. Mobile users are thus allowed to exploit nearby mobile services to boost their computing capabilities without investment in their resource pool. Nevertheless, various challenges, especially its quality-of-service and reliability-aware scheduling, are yet to be addressed. Existing studies and related scheduling strategies consider mobile users to be fully stable and available. In this article, we propose a novel method for reliability-aware and deadline-constrained service composition over opportunistic networks. We leverage the Krill–Herd-based algorithm to yield a deadline-constrained, reliability-aware, and well-executable service composition schedule based on the estimation of completion time and reliability of schedule candidates. We carry out extensive case studies based on some well-known mobile service composition templates and a real-world opportunistic contact data set. The comparison results suggest that the proposed approach outperforms existing ones in terms of success rate and completion time of composed services.Note to Practitioners—Recently, the rapid development of mobile devices and mobile communication leads to the prosperity of mobile service computing. Services running on mobile devices within a limited range are allowed to be composed to coordinate through wireless communication technologies and perform complex tasks and business processes. Despite its great potential, mobile service compositions remains a challenge since the mobility of users and devices imposes high unpredictability on the execution of tasks. A careful investigation into existing methods has found their various limitations, e.g., assuming time-invariant availability of mobile services. This article presents a novel reliability-aware and deadline-constrained service composition method for mobile opportunistic networks. Instead of assuming time-invariant availability of mobile nodes, the proposed method is capable of estimating service availability at run-time and leveraging a Krill–Herd-based algorithm to yield the deadline-constrained, reliability-aware, and well-executable service composition schedules. Case studies based on well-known service composition templates and real-world data sets suggest that it outperforms traditional ones in terms of success and completion time of composed services. It can thus aid the design and optimization of composite services as well as their smooth execution in a mobile environment. It can help practitioners better manage the reliability and performance of real-world applications built upon mobile services.
Qinglan Peng, Yunni Xia, MengChu Zhou, Xin Luo 0001, Yuandou Wang, Chunrong Wu, Mingwei Lin
IEEE Trans Autom. Sci. Eng.6
2020 Decentralized workflow management on software defined infrastructures
abstract
Data-intensive workflow applications are characterized by their continuously growing volumes of data being processing, the complexity of tasks in the pipeline, and infrastructure capacity required for computation and storage. The infrastructure technologies of computing, storage and networking have made tremendous progress during the past yeas. We review the emerging trends in the data-intensive workflow applications, in particular the potential challenges and opportunities enabled by the decentralized application paradigm.
Yuandou Wang, Zhiming Zhao
SERVICES1
2017 Performance Estimation of Fault-prone Infrastructure-as-a-Service Cloud Computing Systems and their Cost-aware Optimal Performance Determination
Kunyin Guo, Yuandou Wang
Mob. Networks Appl.5
2016 On Stochastic Performance and Cost-Aware Optimal Capacity Planning of Unreliable Infrastructure-as-a-Service Cloud
Weiling Li, Yunni Xia, Yuandou Wang, Kunyin Guo, Xin Luo 0001, Mingwei Lin, Wanbo Zheng
ICA3PP4