Zhiming Zhao

dblp:81/2506 · DBLP profile ↗
← Back
112ranked-venue papers
16as first author
60since 2021 · last 2027
0000-0002-6717-9418ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 22 since 2021Software engineering, systems software and programming languages · 29 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 11 since 2021Computer networks · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 HAT-KI: A Hierarchical Auxiliary Task framework with asymmetric Knowledge Injection for early rumor detection
Wenbo Xing, Yan Li 0104, Zhiming Zhao
Inf. Process. Manag.3
2026 A Spatiotemporal Coupling-Based Clustered Federated Learning Scheme for Low Latency Digital Twin Within Heterogeneous IIoT
Miao Liu 0002, Haitao Zhao 0004, Zhiming Zhao, Hongbo Zhu 0002, Dengyin Zhang
IEEE Internet Things J.4
2026 MODIFy : A multi-modal anomaly diagnosis framework with diffusion-enhanced adaptive fusion in microservices
Wujian Zhang, Ruyue Xin, Peng Chen 0007, Ang Bian, Yibin Zhao 0009, Zhiming Zhao
J. Syst. Softw.7
2026 CRL-MM: Context-Aware Relational Learning and Multidimensional Matching for Few-Shot Knowledge Graph Completion
abstract
Few-shot knowledge graph completion (FKGC) aims to infer missing triples for long-tail relationships using a small set of References. Existing FKGC models focus mainly on entity representation aggregation, heavily relying on interactions between central entities and their neighbors. However, real-world knowledge graphs contain relations with multiple semantics, and existing models struggle to capture the diverse semantic information of the relations in different contexts. To address this issue, we propose a novel FKGC model, context-aware relational learning and multidimensional matching (CRL-MM). First, CRL-MM enhances the representation of task relations by obtaining semantic information in different scenarios based on the semantic similarity between task relations and background relations. Second, unlike previous models, which rely mainly on neighborhood relations to capture relation information, CRL-MM considers the entity pair and its neighborhood as a unified contextual whole, aggregating neighborhood information through adaptive task relations and paired entity awareness to improve entity encoding. In addition, during the matching phase, we design a matching network from multiple dimensions, which includes not only the similarity score of the entity pairs but also the triple rationality score to further improve the generalizability of the model. Extensive experiments on public benchmark datasets show that CRL-MM outperforms state-of-the-art methods, and the ablation experiments also demonstrate the effectiveness of each module of the proposed CRL-MM.
Wenchao Jiang, Fangyue Wu, Fanlong Zhang, Quan Chen 0003, Zhiming Zhao, Song Guo 0001
IEEE Trans. Neural Networks Learn. Syst.5
2026 A Hierarchical GNN-Based Multi-Agent Framework for Workflow Scheduling in Hybrid Clouds Considering Privacy Constraints
Hanlin Zhou, Cong Liu 0012, Fang Fang 0007, Zhiming Zhao, Georgios Theodoropoulos 0001, Long Cheng 0003
IEEE Trans. Serv. Comput.5
2025 Unsupervised Detection of Anomalous Commits in Software Repositories
abstract
Identifying anomalous commits is essential for maintaining software quality and reliability, as these anomalies can indicate potential issues in code, development practices, or repository management. Current anomaly detection methods typically rely on prede-fined rules or supervised learning, which suffer from limitations such as dependence on labeled datasets, rigid rule definitions, and high maintenance overhead in rapidly evolving repositories. This paper introduces a novel unsupervised framework for effectively detecting anomalous commits without requiring labeled data or rigid rules, providing a scalable and adaptable solution to enhance code quality in modern version control systems. To address the high-dimensional and mul-tifaceted nature of commit data, our approach com-bines dimensionality reduction techniques with tar-geted feature engineering, enhancing both precision and adaptability in anomaly detection. We systematically evaluate three state-of-the-art unsupervised techniques-Local Outlier Factor (LOF), Isolation Forest (IF), and Histogram-Based Outlier Score (HBOS)-across five diverse open-source repositories. Our results demonstrate that Isolation Forest achieves the highest detection accuracy, effectively balancing precision and recall while capturing both global and local anomalies. Additionally, expert validation confirms the practical relevance of our approach, providing insights into frequent and high-impact anomalies encountered in real-world repositories.
Nafiseh Soveizi, Tomás Candeias, Miroslav Zivkovic, Zhiming Zhao
SSE4
2025 Using Explainable Techniques to Enhance Chain of Thoughts in LLMs
abstract
With the recent advancements in reasoning capabilities with semantic understanding, Large Language Models (LLMs) are increasingly mimicking human reasoning processes and aligning more closely with human values. Yet, unlike how humans sometimes make decisions intuitively by reasoning or understanding the underlying logic, LLMs struggle with transparency in their internal reasoning. Similar to how humans often pause to analyze their reasoning step-by-step to understand their own decisions better, we propose enabling LLMs to perform analogous introspective reasoning using a Chain-of-Thought (CoT) framework powered by SHAP-based explainability. CoT method guides LLMs to explicitly render intermediate reasoning steps, effectively simulating a human’s reflective thought process. Meanwhile, SHAP explainability provides context or logic by quantifying how each semantic token or feature contributes to the reasoning, analogous to a human examining which factors most influenced their choices. Our analysis explores how various training methods influence their semantic comprehension and, consequently, their ability to communicate their reasoning transparently. We apply this approach to a specialized environmental and Earth science ranking dataset—characterized by user feedback and sparse annotations—and we assess whether LLM-generated rankings can not only reflect accurate semantic understanding but also transparently emulate human-like introspection.
Nafis Tanveer Islam, Zhiming Zhao
eScience3
2025 How Good are LLMs at Retrieving Documents in a Specific Domain?
Nafis Tanveer Islam, Zhiming Zhao
FQAS2
2025 Enhancing Adversarial Transferability via Self-Ensemble Feature Alignment
abstract
Deep neural networks (DNNs) have demonstrated remarkable success in tasks such as image classification and object detection, but remain vulnerable to adversarial attacks. To enhance the adversarial transferability across different architectures (e.g., from CNNs to ViTs), existing attacks leverage various strategies such as input transformations, gradient rectification, custom optimization objectives, and model ensembles, but struggle with limited effectiveness under minimal knowledge (e.g., the number of surrogate models). In this work, we propose a novel self-ensemble feature alignment(SEFA) strategy that significantly boosts adversarial transferability with minimal resource overhead. Motivated by the observation that adversarial transferability correlates with feature similarity across models, we leverage Centered Kernel Alignment (CKA) to measure and investigate intermediate features in both inter- and intra-model representation spaces. By splitting a single model into multiple sub-networks and aligning their feature spaces, our method effectively enhances adversarial transferability without relying on additional surrogate models. Experiments on the ImageNet dataset demonstrate that our approach achieves an average ASR of 80.0% (ResNet-50 surrogate) and 75.9% (Inc-v3 surrogate) on ImageNet, which outperform the second-best method (i.e., BSR) by +8.3% and +4.7% respectively. Further, it can seamlessly integrate with existing attacks to further increase cross architecture transferability.
Zhiming Zhao, Qingming Li, Chunyi Zhou 0001, Shouling Ji
ICMR1
2025 D-VRE: From a Jupyter-enabled private research environment to decentralized collaborative research ecosystem
abstract
Today, scientific research is increasingly becoming data-centric and compute-intensive, relying on data and models across distributed sources. However, challenges still exist in the traditional cooperation mode, given the high storage and computing costs, geolocation barriers, and local confidentiality regulations. The Jupyter environment has recently emerged and evolved into a vital virtual research environment for scientific computing, which researchers can use to scale computational analyses up to larger datasets and high-performance computing resources. Nevertheless, existing approaches lack robust support of a decentralized cooperation mode to unlock the full potential of decentralized collaborative scientific research, e.g., seamlessly secure data sharing. In this work, we change the basic structure and legacy norms of current research environments via the seamless integration of Jupyter with Ethereum blockchain capabilities. As such, it creates a Decentralized Virtual Research Environment (D-VRE) from private computational notebooks to a decentralized collaborative research ecosystem. We propose a novel architecture for the D-VRE and prototype some essential D-VRE elements for enabling secure data sharing with decentralized identity, user-centric agreement-making, membership, and research asset management. To validate our method, we conduct an experimental study to test all functionalities of D-VRE smart contracts and their gas consumption. In addition, we deploy the D-VRE prototype on a test net of the Ethereum blockchain for demonstration. The feedback from the studies showcases the current prototype's usability, ease of use, and potential, and suggests further improvements.
Yuandou Wang, Sheejan Tripathi, Siamak Farshidi, Zhiming Zhao
Blockchain Res. Appl.4
2025 Search Multiple Types of Research Assets From Jupyter Notebook
abstract
ABSTRACT Objective Data science and machine learning methodologies are essential to address complex scientific challenges across various domains. These advancements generate numerous research assets such as datasets, software tools, and workflows, which are shared within the open science community. Concurrently, computational notebook environments like Jupyter Notebook, along with platforms like Google Colab and Kaggle Kernel, facilitate data science research and machine learning workflows, transforming data analysis, model development, and knowledge sharing processes. The proliferation of computational notebooks has further enriched the pool of valuable research assets. Researchers frequently require efficient access to these assets to advance their work, yet current tools often require navigating multiple websites and portals, leading to inefficiency and information overload. The challenge is compounded when relying on general web search engines that might not adequately highlight niche scientific resources. Methods To address these issues, we propose the development of an innovative Multiple Research Asset Search (MRAS) system designed to index diverse research assets from heterogeneous sources, offering a unified search interface for researchers. Our system aims to significantly improve the discovery of computational notebooks and datasets, facilitating data‐driven research. Results We developed a pipeline for data extraction and indexing, reviewed and applied state‐of‐the‐art ranking algorithms, enhanced indexing documents with content analysis, and created a Jupyter extension for asset discovery within the working environment. Conclusion This work is structured to detail our approach, literature review, system development, empirical validation, results, and conclusions, illustrating the potential impact of our MRAS system on scientific research efficiency.
Siamak Farshidi, Zhiming Zhao
Softw. Pract. Exp.3
2024 CrowdAL: Towards a Blockchain-empowered Active Learning System in Crowd Data Labeling
abstract
Active Learning (AL) is a machine learning technique where the model selectively queries the most informative data points for labeling by human experts. Integrating AL with crowdsourcing leverages crowd diversity to enhance data labeling but introduces challenges in consensus and privacy. This poster presents CrowdAL, a blockchain-empowered crowd AL system designed to address these challenges. CrowdAL integrates blockchain for transparency and a tamper-proof incentive mechanism, using smart contracts to evaluate crowd workers’ performance and aggregate labeling results, and employs zeroknowledge proofs to protect worker privacy.
Shaojie Hou, Yuandou Wang, Zhiming Zhao
e-Science3
2024 A Collaborative Framework for Facilitating Federated Learning among Jupyter Users
abstract
Federated learning (FL) allows multiple partners to train machine learning models without sharing raw data, thus preserving privacy. Despite its promising aspects, existing FL frameworks have some drawbacks regarding flexibility, decentralized aggregation, and collaborative environments. This poster presents FedLearn, a collaborative community framework built atop the JupyterLab environment for FL among Jupyter users. We use a microservices architecture to implement the framework and enable automated FL deployment processes across multiple clouds. The demonstration showcases the feasibility of the FedLearn community portal for Jupyter users.
Anandan Krishnasamy, Yuandou Wang, Zhiming Zhao
e-Science3
2024 PriCE: Privacy-Preserving and Cost-Effective Scheduling for Parallelizing the Large Medical Image Processing Workflow over Hybrid Clouds
Yuandou Wang, Neel Kanwal, Kjersti Engan, Chunming Rong, Paola Grosso, Zhiming Zhao
Euro-Par (1)6
2024 MARS: Multi-Agent Deep Reinforcement Learning for Real-Time Workflow Scheduling in Hybrid Clouds with Privacy Protection
abstract
Scheduling workflows in hybrid cloud environments presents significant challenges due to the inherent complexity of workflows and the dynamic nature of cloud resources. This complexity is further increased when attempting to balance workflow performance with privacy protection. Recent efforts have leveraged deep reinforcement learning (DRL) to address these challenges. However, most of these approaches rely on single-agent models, which can lead to security issues and scalability problems due to their centralized processing. Specifically, the properties of workflows are transferred to the single agent, which risks leaking privacy information. Our paper addresses these issues by introducing MARS, a real-time workflow scheduling method that prioritizes privacy protection in hybrid clouds. MARS leverages multi-agent deep reinforcement learning (MADRL) to optimize the workflow scheduling of cloud virtual machines (VMs). The benefit of our solution is that it relies on the collaborative learning of multi-agents on multiple VMs, which could assign user data to specific cloud servers for privacy protection while sharing training experiences between agents. In our implementation, MARS aims to reduce workflow completion time and operational costs while complying with strict privacy protection guidelines. The experimental results demonstrate that MARS can significantly surpass existing methods, reducing makespan by an average of $53.18 \%$ and costs by $61.98 \%$ compared to basic techniques, and achieving $20.26 \%$ and $25.71 \%$ improvements over the latest advanced methods, respectively.
Long Cheng 0003, Haoyang He, Qingzhi Liu, Zhiming Zhao, Fang Fang 0007
ICPADS5
2024 Preface of special issue on Artificial Intelligence for time-critical computing systems
Long Cheng 0003, Zhiming Zhao
Future Gener. Comput. Syst.3
2024 Autonomous selection of the fault classification models for diagnosing microservice applications
Yujia Song, Ruyue Xin, Peng Chen 0007, Rui Zhang 0099, Zhiming Zhao
Future Gener. Comput. Syst.6
2024 A fine-grained robust performance diagnosis framework for run-time cloud applications
abstract
To maintain the required service quality of time-critical cloud applications, operators must continuously monitor their runtime status, detect potential performance anomalies, and diagnose the root causes of these anomalies effectively. However, existing performance diagnosis methods face challenges such as the need for high-quality labeled data, the low reusability and robustness of performance anomaly detection models, and the absence of real-time fine-grained root cause localization. These challenges make fixing performance issues quickly and developing effective adaptation decisions difficult. We provide a Fine-grained Robust Performance Diagnosis (FIRED) framework to tackle those challenges. The framework offers a metrics selection component to filter noise and improve detection efficiency, an anomaly detection component that assembles several well-selected base models with a deep neural network, and adopts weakly supervised learning considering fewer labels exist in reality. The framework also employs a real-time, fine-grained root cause localization component to locate dependent resource metrics of performance anomalies. Our experiments show that the framework can effectively reduce data noise and achieve the best accuracy and algorithm robustness for performance anomaly detection. In addition, the framework can accurately localize the first root causes, with an average accuracy higher than 0.7 for locating the first four root cause metrics.
Ruyue Xin, Peng Chen 0007, Paola Grosso, Zhiming Zhao
Future Gener. Comput. Syst.4
2024 HT-RCM: Hashimoto's Thyroiditis Ultrasound Image Classification Model Based on Res-FCT and Res-CAM
abstract
The early lesions of Hashimoto's thyroiditis are inconspicuous, and the ultrasonic features of these early lesions are indistinguishable from other thyroid diseases. This paper proposes a Hashimoto Thyroiditis ultrasound image classification model HT-RCM which consists of a Residual Full Convolution Transformer (Res-FCT) model and a Residual Channel Attention Module (Res-CAM). To collect the low-order information caused by hypoechoic signals accurately, the residual connection is injected between FCTs to form Res-FCT which helps HT-RCM superimpose the low-order input information and high-order output information together. Res-FCT can make HT-RCM focus more on hypoechoic information while avoiding gradient dispersion. The initial feature map is inserted into Res-FCT again through a down-sampling component, which further helps HT-RCM exact multi-level original semantic information in the ultrasound image. Res-CAM is constructed by implementing a residual connection between a channel attention module and a convolution layer. Res-CAM can effectively increase the weights of the lesion channels while suppressing the weights of the noise channels, which makes HT-RCM focus more on the lesion regions. The experimental results on our collected dataset show that HT-RCM outperforms the mainstream models and obtains state-of-the-art performance in HT ultrasound image classification.
Wenchao Jiang, Tianchun Luo, Guanghui Yue 0001, Zhiming Zhao, Jianxuan Wen
IEEE J. Biomed. Health Informatics6
2024 FBENet: Feature-Level Boosting Ensemble Network for Hashimoto's Thyroiditis Ultrasound Image Classification
abstract
Distinguishing Hashimoto's thyroiditis (HT) lesions from ordinary thyroid tissues is difficult with ultrasound images. Challenges in achieving high performance of HT ultrasound image classification include the low resolution, blurred features and large area of irrelevant noise. To address these problems, we propose a Feature-level Boosting Ensemble Network (FBENet) for HT ultrasound image classification. Specifically, to capture the features of suspicious HT lesions efficiently, an Ensemble Feature Boosting Module (EFBM) is introduced into the feature-level ensemble to boost the blurred features. Then, the spatial attention mechanism is adopted in backbone models to improve the feature focusing performance and representation ability. Furthermore, feature-level ensemble technique is employed in the training process to achieve more comprehensive feature representation ability. Experimentally, FBENet was trained on 6,503 HT ultrasound images, and tested on 1,626 HT ultrasound images with 82.92% accuracy and 89.24% AUC on average.
Wenchao Jiang, Tianchun Luo, Ji He 0001, Zhiming Zhao, Jianxuan Wen
IEEE J. Biomed. Health Informatics6
2024 SurgNet: Self-Supervised Pretraining With Semantic Consistency for Vessel and Instrument Segmentation in Surgical Images
abstract
Blood vessel and surgical instrument segmentation is a fundamental technique for robot-assisted surgical navigation. Despite the significant progress in natural image segmentation, surgical image-based vessel and instrument segmentation are rarely studied. In this work, we propose a novel self-supervised pretraining method (SurgNet) that can effectively learn representative vessel and instrument features from unlabeled surgical images. As a result, it allows for precise and efficient segmentation of vessels and instruments with only a small amount of labeled data. Specifically, we first construct a region adjacency graph (RAG) based on local semantic consistency in unlabeled surgical images and use it as a self-supervision signal for pseudo-mask segmentation. We then use the pseudo-mask to perform guided masked image modeling (GMIM) to learn representations that integrate structural information of intraoperative objectives more effectively. Our pretrained model, paired with various segmentation methods, can be applied to perform vessel and instrument segmentation accurately using limited labeled data for fine-tuning. We build an Intraoperative Vessel and Instrument Segmentation (IVIS) dataset, comprised of ~3 million unlabeled images and over 4,000 labeled images with manual vessel and instrument annotations to evaluate the effectiveness of our self-supervised pretraining method. We also evaluated the generalizability of our method to similar tasks using two public datasets. The results demonstrate that our approach outperforms the current state-of-the-art (SOTA) self-supervised representation learning methods in various surgical image segmentation tasks.
Hu Han 0001, Zhiming Zhao, Xilin Chen 0001
IEEE Trans. Medical Imaging4
2024 LHNetV2: A Balanced Low-Cost Hybrid Network for Single Image Dehazing
abstract
Single-image dehazing is a challenging task that requires both local details and global distribution. Existing methods face challenges in color imbalance and inconsistent details when predicting a haze-free image, because of their limitations in generalization from a specific setting (physics-based methods), capturing global information (CNN-based methods) and capturing detailed local information ( ViT-based methods). In response to these challenges, we propose a balanced low-cost hybrid network called LHNetV2 based on LHNetV1. The key insight of LHNetV2 is the effective fusion of different features, and a series of novel approaches is proposed to increase the running speed of the original LHNetV1. Firstly, building upon the Feature-aware Information Fusion method, we preserve the original Physical Embedding and Architecture Aggregation components in LHNetV1. Next, to overcome the speed bottleneck of LHNetV1, we enhance the calculation method of attention in the ViT sub-network and streamline the cross-stage interaction strategy in the CNN main-network. Finally, we introduce a dynamic adversarial loss function to bolster both the training stability and performance of LHNetV2. The experiments are extensively conducted on mainstream datasets, and the results demonstrate that LHNetV2 achieves the best balance between the performance and the running speed in single-image dehazing. The code is available at https://github.com/SHYuanBest/LHNet.
Shenghai Yuan 0002, Jijia Chen, Wenchao Jiang, Zhiming Zhao, Song Guo 0001
IEEE Trans. Multim.4
2024 A Deep Reinforcement Learning-Based Preemptive Approach for Cost-Aware Cloud Job Scheduling
abstract
With some specific characteristics such as elastics and scalability, cloud computing has become the most promising technology for online business nowadays. However, how to efficiently perform real-time job scheduling in cloud still poses significant challenges. The reason is that those jobs are highly dynamic and complex, and it is always hard to allocate them to computing resources in an optimal way, such as to meet the requirements from both service providers and users. In recent years, various works demonstrate that deep reinforcement learning (DRL) can handle real-time cloud jobs well in scheduling. However, to our knowledge, none of them has ever considered extra optimization opportunities for the allocated jobs in their scheduling frameworks. Given this fact, in this work, we introduce a novel DRL-based preemptive method for further improve the performance of the current studies. Specifically, we try to improve the training of scheduling policy with effective job preemptive mechanisms, and on that basis to optimize job execution cost while meeting users' expected response time. We introduce the detailed design of our method, and our evaluations demonstrate that our approach can achieve better performance than other scheduling algorithms under different real-time workloads, including the DRL approach.
Long Cheng 0003, Yue Wang 0073, Cheng Liu 0008, Zhiming Zhao, Ying Wang 0001
IEEE Trans. Sustain. Comput.5
2023 Ocean Data Quality Assessment through Outlier Detection-enhanced Active Learning
abstract
Ocean and climate research benefits from global ocean observation initiatives such as Argo, GLOSS, and EMSO. The Argo network, dedicated to ocean profiling, generates a vast volume of observatory data. However, data quality issues from sensor malfunctions and transmission errors necessitate stringent quality assessment. Existing methods, including machine learning, fall short due to limited labeled data and imbalanced datasets. To address these challenges, we propose an Outlier Detection-Enhanced Active Learning (ODEAL) framework for ocean data quality assessment, employing Active Learning (AL) to reduce human experts’ workload in the quality assessment workflow and leveraging outlier detection algorithms for effective model initialization. We also conduct extensive experiments on five large-scale realistic Argo datasets to gain insights into our proposed method, including the effectiveness of AL query strategies and the initial set construction approach. The results suggest that our framework enhances quality assessment efficiency by up to 465.5% with the uncertainty-based query strategy compared to random sampling and minimizes overall annotation costs by up to 76.9% using the initial set built with outlier detectors.
Yiyang Qi, Ruyue Xin, Zhiming Zhao
IEEE Big Data4
2023 Decoding NFT Market Dynamics with Data
abstract
Non-Fungible Tokens (NFTs) are a developing area in the market of digital assets. NFTs represent digital or real-world items like artwork, gaming collectibles and real estate. We aim to study the daily working of NFT market and its interaction with cryptocurrency (Ether and Bitcoin) and search interest.Our approach involves identification of models encompassing both global and local feature importance. Various regression methods are utilized to determine the feature importance and select the predictive features effectively. Moreover, this study explores the relationship between search interest and weekly NFT sales and vice versa, to comprehend how public interest impacts the NFT market. Lastly, anomalies in daily sales are detected and analysed using STL Decomposition and SHAPely.The study reveals that intrinsic sales attributes and trade profits drive daily NFT sales, with positive sentiment significantly impacting Ethereum volatility and NFT sales. External factors like NFT supply, Ether price, and trade profits also influence anomalies. Positive sentiment significantly shapes crypto and NFT market dynamics.
Smruti Inamdar, Zhiming Zhao
CloudCom2
2023 Increasing Robustness of Blockchain Peer-to-Peer Networks with Alternative Peer Initialization
abstract
In this paper, we identified the most important vulnerabilities of a peer-to-peer (P2P) network in a blockchain and proposed an algorithm for adding new peers to an existing network in order to increase robustness. We determined the main vulnerabilities and their detection metrics by means of current literature. They can be categorized by topological characteristics, network resource statistics, the local characteristics of the neighbors set per peer, and implementation details of various parts of the blockchain. Based on topological metrics, we set forth an algorithm for adding new peers into an optimal position in order to increase overall robustness. Her we define robustness as the functionality and stability of the blockchain. A proof of concept, comparing it to random connections showed no indication of a better robustness, as measured by chain creation and its stability, over time. However future works is needed to determine efficacy of this algorithm and other ways of using metrics to initialize peer connections.
Bernadet Klein Wassink, Zhiming Zhao
CloudCom2
2023 Towards a Service-based Adaptable Data Layer for Cloud Workflows
abstract
Many scientific workflows are data-driven and need to be continuously executed for the large volume of datasets transferred from distributed data sources. The overhead arising from data transfers must be considered when optimizing workflow performance. Many workflow systems support various data transfer protocols (DTPs) and file systems. However, challenges that hinder wide protocol adoption are mainly the need for more feasibility of adapting new solutions, such as decentralized ones. In this paper, we prototype a container-native data layer that supports multiple DTPs, e.g., FTP, WebDAV, and IPFS, for Cloud workflows. Based on this tool, we demonstrated the feasibility of using combinations of Docker, CWL, and Argo to deploy and execute several application scenarios adaptably. Besides, we analyzed the performance of data transfers and workflow execution time between IPFS and WebDAV, which can help users decide which one to handle data. Our results show that IPFS outperforms WebDAV in uploading large files, and the makespan via IPFS executed in Argo is comparable with WebDAV.
Yuandou Wang, Nikita Janse, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
COMPSAC5
2023 Towards a Knowledge Graph Enhanced Automation and Collaboration Framework for Digital Twins
abstract
The Digital Twin (DT) provides a digital representation of a physical system and allows users to interactively study the physical processes of a real system via the digital representation in different scenarios in real time. The development of a DT is highly complex; it requires not only expertise from multiple disciplines but also the integration of often heterogeneous software components, e.g., simulations, machine learning, visualization, and user interface components across distributed environments. This poster presents a Knowledge Graph-based ontological framework to boost automation and collaboration during the DT lifecycle stages. We implement our methods in developing a what-if analysis service for a DT of an ecosystem of wetlands and its automated deployment to the Amazon Web Services (AWS) cloud.
Vasileios Christou, Yuandou Wang, Zhiming Zhao
e-Science3
2023 CWL-FLOps: A Novel Method for Federated Learning Operations at Scale
abstract
Federated Learning (FL) has attracted much attention in recent years because it enables users with private data sets to train a global model collaboratively without raw data exchange. However, due to a lack of automation, researchers often struggled to develop, deploy, track, and manage all the data, steps, and configuration setup for all FL participating nodes. Federated Learning Operations (FLOps) is recently emerging in the FL community, a new methodology for developing FL systems efficiently and continuously. Some research works discussed approaches for FLOps, but only a few solutions address managing FL application scenarios from the workflow perspective. This poster proposes CWL-FLOps, a novel CWL-based method for FLOps, which can improve the flexibility of FL abstraction and fully automate the FL deployment and execution by mapping high-level descriptions onto distributed resource nodes. Our experiments demonstrate the feasibility of describing centralized and decentralized FL scenarios using CWL abstracted definitions without relying on heavily customized or external software for execution.
Chronis Kontomaris, Yuandou Wang, Zhiming Zhao
e-Science3
2023 A Dense Retrieval System and Evaluation Dataset for Scientific Computational Notebooks
abstract
The discovery and reutilization of scientific codes are crucial in many research activities. Computational notebooks have emerged as a particularly effective medium for sharing and reusing scientific codes. Nevertheless, effectively locating relevant computational notebooks is a significant challenge. First, computational notebooks encompass multi-modal data comprising unstructured text, source code, and other media, posing complexities in representing such data for retrieval purposes. Second, the absence of evaluation datasets for the computational notebook search task hampers fair performance assessments within the research community. Prior studies have either treated computational notebook search as a code-snippet search problem or focused solely on content-based approaches for searching computational notebooks. To address the aforementioned difficulties, we present DeCNR, tackling the information needs of researchers in seeking computational notebooks. Our approach leverages a fused sparse-dense retrieval model to represent computational notebooks effectively. Additionally, we construct an evaluation dataset including actual scientific queries, computational notebooks, and relevance judgments for fair and objective performance assessment. Experimental results demonstrate that the proposed method surpasses baseline approaches in terms of F1@5 and NDCG@5. The proposed system has been implemented as a web service shipped with REST APIs, allowing seamless integration with other applications and web services.
Zhiming Zhao
e-Science3
2023 Integrating R in a Distributed Scientific Workflow via a Jupyter-Based Environment
abstract
The Research Infrastructure Lifewatch Italy has developed a Virtual Research Environment for studies on phytoplankton ecology that includes computational services based on R, a programming language widely used for data science and ecology. Here we have verified the feasibility of a Jupyter-based research environment, the NaaVRE, which has so far been tested only with Python, for running R code in a workflow on the Cloud. The successful execution demonstrated the potentialities of R in Cloud-based research environments. However, further investigation is needed, in particular, to overcome the issue of the lack of dependencies declaration in R. The possibility of performing analyses in a workflow, combined with the computational resources of remote infrastructures, will support scientists in carrying out FAIR and innovative research in a more efficient, integrated and collaborative way.
Mariantonietta La Marra, Diederik Blanson Henkemans, Jessica Titocci, Spiros Koulouzis, Ilaria Rosati, Zhiming Zhao
e-Science6
2023 A Survey on Dataset Distillation: Approaches, Applications and Future Directions
abstract
Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, including support for continual learning, neural architecture search, and privacy protection. Despite recent advances, we lack a holistic understanding of the approaches and applications. Our survey aims to bridge this gap by first proposing a taxonomy of dataset distillation, characterizing existing approaches, and then systematically reviewing the data modalities, and related applications. In addition, we summarize the challenges and discuss future directions for this field of research.
Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong
IJCAI7
2023 Cost-aware scheduling systems for real-time workflows in cloud: An approach based on Genetic Algorithm and Deep Reinforcement Learning
Long Cheng 0003, Cong Liu 0012, Zhiming Zhao, Ying Mao 0001
Expert Syst. Appl.4
2023 Robustness challenges in Reinforcement Learning based time-critical cloud resource scheduling: A Meta-Learning based solution
abstract
Cloud computing attracts increasing attention in processing dynamic computing tasks and automating the software development and operation pipeline. In many cases, the computing tasks have strict deadlines. The cloud resource manager (e.g., orchestrator) effectively manages the resources and provides tasks Quality of Service (QoS). Cloud task scheduling is tricky due to the dynamic nature of task workload and resource availability. Reinforcement Learning (RL) has attracted lots of research attention in scheduling. However, those RL-based approaches suffer from low scheduling performance robustness when the task workload and resource availability change, particularly when handling time-critical tasks. This paper focuses on both challenges of robustness and deadline guarantee among such RL, specifically Deep RL (DRL)-based scheduling approaches. We quantify the robustness measurements as the retraining time and investigate how to improve both robustness and deadline guarantee of DRL-based scheduling. We propose MLR-TC-DRLS, a practical, robust Meta Deep Reinforcement Learning-based scheduling solution to provide time-critical tasks deadline guarantee and fast adaptation under highly dynamic situations. We comprehensively evaluate MLR-TC-DRLS performance against RL-based and RL advanced variants-based scheduling approaches using real-world and synthetic data. The evaluations validate that our proposed approach improves the scheduling performance robustness of typical DRL variants scheduling approaches with 97%–98.5% deadline guarantees and 200%–500% faster adaptation.
Hongyun Liu, Peng Chen 0007, Xue Ouyang 0003, Hui Gao 0003, Bing Yan 0001, Paola Grosso, Zhiming Zhao
Future Gener. Comput. Syst.7
2023 Identifying performance anomalies in fluctuating cloud environments: A robust correlative-GNN-based explainable approach
Yujia Song, Ruyue Xin, Peng Chen 0007, Rui Zhang 0099, Zhiming Zhao
Future Gener. Comput. Syst.6
2023 CausalRCA: Causal inference based precise fine-grained root cause localization for microservice applications
abstract
Effectively localizing root causes of performance anomalies is crucial to enabling the rapid recovery and loss mitigation of microservice applications in the cloud. Depending on the granularity of the causes that can be localized, a service operator may take different actions, e.g., restarting or migrating services if only faulty services can be localized (namely, coarse-grained) or scaling resources if specific indicative metrics on the faulty service can be localized (namely, fine-grained). Prior research mainly focuses on coarse-grained faulty service localization, and there is now a growing interest in fine-grained root cause localization to identify faulty services and metrics. Causal inference (CI) based methods have gained popularity recently for root cause localization, but currently used CI methods have limitations, such as the linear causal relations assumption and strict data distribution requirements. To tackle these challenges, we propose a framework named CausalRCA to implement fine-grained, automated, and real-time root cause localization. The CausalRCA uses a gradient-based causal structure learning method to generate weighted causal graphs and a root cause inference method to localize root cause metrics. We conduct coarse- and fine-grained root cause localization to evaluate the localization performance of CausalRCA. Experimental results show that CausalRCA has significantly outperformed baseline methods in localization accuracy, e.g., the average AC@3 of the fine-grained root cause metric localization in the faulty service is 0.719, and the average increase is 10% compared with baseline methods. In addition, the average Avg@5 has improved by 9.43%. Codes and data are open-sourced and can be found in our Github repository CausalRCA.
Ruyue Xin, Peng Chen 0007, Zhiming Zhao
J. Syst. Softw.3
2023 Understanding the Necessity and Economic Benefits of Lockdown Measures to Contain COVID-19
abstract
Since the outbreak of the coronavirus disease 2019 (COVID-19), the issue of how to maintain economic development while containing the epidemic has become a significant concern for decision-makers. Though lockdown measures are verified to be very effective in containing the epidemic, its economic costs and other influences have not been fully explored. As a result, decision-makers in many countries are still hesitant to include the lockdown measure in an intervention strategy in response to COVID-19. To address this issue, we propose a universal computational experiment approach for policy evaluation and adjustment based on the Artificial societies, Computational experiments, Parallel execution (ACP) concept. First, we innovatively construct a model via observable CO2 emissions, which is able to estimate the economic costs affected by nonpharmaceutical interventions. Furthermore, based on the population movement data, a risk source model is proposed to estimate the local transmission risk for any prefectures outside the epicenter. Finally, we integrate the data models in a high-resolution agent-based artificial society and carry out large-scale computational experiments supported by the Tianhe supercomputer. Policy adjustments and evaluations are carried out in four cities: Wenzhou, Guangzhou, Beijing, and Wuhan. Our research findings show important implications for policy-making: 1) the local transmission of a city can be almost contained if lockdowns are adopted immediately when the risk index is larger than 1.645, 1.960, or 2.576 at the 90%, 95%, or 99% confidence interval, respectively; 2) if lockdowns are required, in-advance lockdown measures facilitate mitigation efficacy and reduce economic loss; and 3) lockdowns lasting for 7–14 days in a prefecture would be effective in controlling the spread of the epidemic. The duration of the measure should be prolonged with the increment of the initial transmission risk.
Zhengqiu Zhu, Chuan Ai, Bin Chen 0003, Wei Duan 0002, Xiaogang Qiu, Xin Lu 0002, Zhiming Zhao, Zhong Liu 0002
IEEE Trans. Comput. Soc. Syst.9
2022 Multi-Objective Robust Workflow Offloading in Edge-to-Cloud Continuum
abstract
Workflow offloading in the edge-to-cloud continuum copes with an extended calculation network among edge devices and cloud platforms. With the growing significance of edge and cloud technologies, workflow offloading among these environments has been investigated in recent years. However, the dynamics of offloading optimization objectives, i.e., latency, resource utilization rate, and energy consumption among the edge and cloud sides, have hardly been researched. Consequently, the Quality of Service(QoS) and offloading performance also experience uncertain deviation. In this work, we propose a multi-objective robust offloading algorithm to address this issue, dealing with dynamics and multi-objective optimization. The workflow request model in this work is modeled as Directed Acyclic Graph(DAG). An LSTM-based sequence-to-sequence neural network learns the offloading policy. We then conduct comprehensive implementations to validate the robustness of our algorithm. As a result, our algorithm achieves better offloading performance regarding each objective and faster adaptation to newly changed environments than fine-tuned typical single-objective RL-based offloading methods.
Hongyun Liu, Ruyue Xin, Peng Chen 0007, Zhiming Zhao
CLOUD4
2022 MOGPlay: A Decentralized Crowd Journalism Application for Democratic News Production
abstract
Media production and consumption behaviors are changing in response to new technologies and demands, giving birth to a new generation of social applications. Among them, crowd journalism represents a novel way of constructing democratic and trustworthy news relying on ordinary citizens arriving at breaking news locations and capturing relevant videos using their smartphones. The ARTICONF project [1] proposes a trustworthy, resilient, and globally sustainable toolset for developing decentralized applications (DApps). Leveraging the ARTICONF tools, we introduce a new DApp for crowd journalism called MOGPlay. MOGPlay collects and manages audio-visual content generated by citizens and provides a secure blockchain platform that rewards all stakeholders involved in professional news production. Besides live streaming, MOGPlay offers a marketplace for audio-visual content trading among citizens and free journalists with an internal token ecosystem. We discuss the functionality and implementation of the MOGPlay DApp and illustrate three pilot crowd journalism live scenarios that validate the prototype.
Inês Rito Lima, Cláudia Marinho, Vasco Filipe, Alexandre Ulisses, Nishant Saurabh, Antorweep Chakravorty, Zhiming Zhao, Atanas Hristov, Radu Prodan
ASONAM7
2022 Context-Aware Notebook Search in a Jupyter-Based Virtual Research Environment
abstract
Computational notebook environments such as the Jupyter play an increasingly important role in data-centric research for prototyping computational experiments, documenting code implementations, and sharing scientific results. Effectively discovering and reusing notebooks available on the web can reduce repetitive work and facilitate scientific innovations. However, general-purpose web search engines (e.g., Google Search) do not explicitly index the contents of notebooks, and notebook repositories (e.g., Kaggle and GitHub) require users to create domain-specific queries based on the metadata in the notebook catalogs, which fail to capture the working contexts in the notebook environment. This poster presents a Context-aware Notebook Search Framework (CANSF) to enable a researcher to seamlessly discover external notebooks based on semantic contexts of the literate programming activities in the Jupyter environment.
Siamak Farshidi, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
e-Science5
2022 The Extreme Counts: Modeling the Performance Uncertainty of Cloud Resources with Extreme Value Theory
Mengjuan Li, Jinshu Su, Hongyun Liu, Zhiming Zhao, Xue Ouyang 0003, Huan Zhou 0006
ICSOC4
2022 Federating Unlabeled Samples: A Semi-supervised Collaborative Framework for Whole Slide Image Analysis
Laëtitia Launet, Rocío del Amor, Adrián Colomer, Andrés Mosquera-Zamudio, Anaïs Moscardó, Carlos Monteagudo Mañas, Zhiming Zhao, Valery Naranjo
IDEAL7
2022 An Adaptable Indexing Pipeline for Enriching Meta Information of Datasets from Heterogeneous Repositories
Siamak Farshidi, Zhiming Zhao
PAKDD (2)2
2022 Effectively Detecting Operational Anomalies In Large-Scale IoT Data Infrastructures By Using A GAN-Based Predictive Model
abstract
Abstract Quality of data services is crucial for operational large-scale internet-of-things (IoT) research data infrastructure, in particular when serving large amounts of distributed users. Effectively detecting runtime anomalies and diagnosing their root cause helps to defend against adversarial attacks, thereby essentially boosting system security and robustness of the IoT infrastructure services. However, conventional anomaly detection methods are inadequate when facing the dynamic complexities of these systems. In contrast, supervised machine learning methods are unable to exploit large amounts of data due to the unavailability of labeled data. This paper leverages popular GAN-based generative models and end-to-end one-class classification to improve unsupervised anomaly detection. A novel heterogeneous BiGAN-based anomaly detection model Heterogeneous Temporal Anomaly-reconstruction GAN (HTA-GAN) is proposed to make better use of a one-class classifier and a novel anomaly scoring function. The Generator-Encoder-Discriminator BiGAN structure can lead to practical anomaly score computation and temporal feature capturing. We empirically compare the proposed approach with several state-of-the-art anomaly detection methods on real-world datasets, anomaly benchmarks and synthetic datasets. The results show that HTA-GAN outperforms its competitors and demonstrates better robustness.
Peng Chen 0007, Hongyun Liu, Ruyue Xin, Thierry Carval, Yunni Xia, Zhiming Zhao
Comput. J.7
2022 A Bayesian game-enhanced auction model for federated cloud services using blockchain
abstract
Industrial applications often require federated cloud services from multiple providers to improve reliability and flexibility. Traditional selection methods through auctions usually involve a centralized auctioneer to coordinate the auction procedure. Blockchain and smart contracts provide a decentralized mechanism to automate the cloud auction process; however, existing solutions fail in the selection of the most suitable providers and the violation detection of the signed auction agreements, which are also known as service-level agreements (SLAs). To tackle these problems, we propose an integrated auction model using Bayesian game theory and blockchain techniques. The proposed model is enhanced with two Bayesian Nash Equilibriums (BNEs); the first BNE enables the selection of cost-effective providers to construct the federated cloud services, while the second BNE ensures consistent and trustworthy monitoring of federated SLAs. Moreover, a timed message submission (TMS) algorithm is proposed to protect the auction privacy during the message submission phase. This paper validates the equilibrium results of two BNEs and implements the proposed model on the Ethereum blockchain. The analytical and experimental results demonstrate the feasibility, trustworthiness, and cost-effectiveness of our model.
Zeshun Shi, Huan Zhou 0006, Cees T. A. M. de Laat, Zhiming Zhao
Future Gener. Comput. Syst.4
2022 Label entropy-based cooperative particle swarm optimization algorithm for dynamic overlapping community detection in complex networks
abstract
The real-world complex networks, such as biological, transportation, biomedical, web, and social networks, are usually dynamic and change over time. The communities which reflect the substructures hidden in the networks usually overlap each other, and detecting overlapping communities in the dynamic complex networks is a challenging task. Prior researchers have applied multiobjective optimization method to the detection of dynamic overlapping communities and achieved some excellent results. However, in terms of multiobjective processing, the prior studies all adopt the decomposition method based on weight parameters, and different weight parameters or different parameter values can easily affect the community detection results which further results in the uneven distribution of the detected results in the target space. To solve the above problems, a hybrid algorithm, that is, Collaborative Particle Swarm multiobjective Optimization-based Dynamic Overlapping Community Detection (CPSO-DOCD) algorithm is proposed in this paper. First, to improve the diversity of particles, the encoding/decoding of the particle and the cross inheritance and the variation of particle are redefined first based on label propagation. In each network snapshot, multiple particle swarms are initialized based on Community Overlap Propagation Algorithm (COPRA) to generate particles with uniform distribution. Multiple different objective functions are optimized using multiple particle swarms respectively to avoid the incorrect selection of weight parameters. In addition, a reference-point-based is adopted in the particle selecting stage to solve the uneven distribution of detected results in the target space. Second, a node label entropy-based particle swarm algorithm is proposed to improve the accuracy of community detection of current network snapshots. Finally, when one snapshot switches to another over time, a migration strategy based on COPRA local-search and clique generation is utilized to adjust the prior community detection results, which enables the former results can be adapted to the new network snapshots. The experiments are implemented based on four dynamic networks which are Cit-HepPh, Cit-HepTh, Emailed-EU-core-temporal, and CollegeMsg. The hypervolume value of the overlapping community detection result obtained by CPSO-DOCD is 0.5%–2% higher than MDOA, MCMOEA, SLPAD, and iLCD. Furthermore, CPSO-DOCD also performed better than MDOA, MCMOEA, SLPAD, and iLCD on C-metric values, and CPSO-DOCD can approach approximately to the Pareto frontier.
Wenchao Jiang, Shucan Pan, Chaohai Lu, Zhiming Zhao, Sui Lin, Meng Xiong, Zhongtang He
Int. J. Intell. Syst.4
2022 Metric learning-based whole health indicator model for industrial robots
abstract
Aiming at the problems of complex structure, high components coupling, and difficultly monitoring of the whole health status with the industrial robot, a metric learning-based whole health indicator model is proposed. First, according to the more obvious degradation characteristics of industrial robots during accelerated operation, the accelerated signal is segmented and then the time-domain features are extracted. Second, the long-term and short-term memory (LSTM) network combined with the multihead attention is used to construct the network model, and the metric learning method is adopted to learn the similarity measurement method of the industrial robot monitoring data. Finally, the similarity measure method got from metric learning is used to construct the whole health indicator, which describes the whole degradation trend of the industrial robot. The experiments are based on the real accelerated aging data set from industrial robots. The results show that the proposed model can effectively construct the whole health indicator for industrial robots. The average trend of the proposed model reaches 0.9769. The average monotonicity reaches 0.5666, which is 0.1748, 0.1577, and 0.1492 higher than the similarity measurement method based on Euclidean distance, Markov distance, and LSTM.
Ping Li 0045, Hanlin Zeng, Tiancai Liang, Wenchao Jiang, Zhiming Zhao
Int. J. Intell. Syst.6
2022 Real-time recognition and warning of mask wearing based on improved YOLOv5 R6.1
abstract
Since the new crown epidemic, mask-wearing has become a new normal in people's work and life. The inspection mechanism for mask-wearing at the entrance and exit of public places is seriously insufficient. The phenomenon of “pick-up on entry” has led to the severe formalization of mask-wearing inspection. Manual detection of mask-wearing in an open and dynamic crowded environment is unrealistic, which is not only time-consuming and labor-intensive but also cannot achieve early warning throughout the entire process. In response to this problem, this paper proposes a real-time recognition and early warning method for mask-wearing in an open, dynamic, complex environment based on improved YOLOv5 R6.1. First, replacing the first Conv structure of the backbone network in the YOLOv5 R6.1 model with an improved Stem structure to minimize the computational overhead while improving the performance. Then by normalizing the data, the random erasure data expansion technique is used to enhance the antiocclusion robustness of the algorithm. Finally, according to the mask-wearing specification in the training data set, optimizing and adjusting the anchor box parameters of the YOLOv5 R6.1 model to improve the model's ability to recognize small targets. The experiments are based on open data sets, and the results show that the mean precision (mAP), precision, and recall of this method reach 92.9%, 94.1%, and 88.5% on average, and the average frames per second (FPS) reaches 117. Moreover, the mAP and FPS are improved by an average of 6.5% and 474% compared with algorithms based on RetinaNet, Attention-Retina, Single Shot multibox Detector, Fast-RCNN, YOLOv4, and YOLOv5.
Shenghai Yuan 0001, Tiancai Liang, Wenchao Jiang, Sui Lin, Zhiming Zhao
Int. J. Intell. Syst.6
2022 Short-text feature expansion and classification based on nonnegative matrix factorization
abstract
In this paper, a non-negative matrix factorization feature expansion (NMFFE) approach was proposed to overcome the feature-sparsity issue when expanding features of short-text. First, we took the internal relationships of short texts and words into account when segmenting words from texts and constructing their relationship matrix. Second, we utilized the Dual regularization non-negative matrix tri-factorization (DNMTF) algorithm to obtain the words clustering indicator matrix, which was used to get the feature space by dimensionality reduction methods. Thirdly, words with close relationship were selected out from the feature space and added into the short-text to solve the sparsity issue. The experimental results showed that the accuracy of short text classification of our NMFFE algorithm increased 25.77%, 10.89%, and 1.79% on three data sets: Web snippets, Twitter sports, and AGnews, respectively compared with the Word2Vec algorithm and Char-CNN algorithm. It indicated that the NMFFE algorithm was better than the BOW algorithm and the Char-CNN algorithm in terms of classification accuracy and algorithm robustness.
Wenchao Jiang, Zhiming Zhao
Int. J. Intell. Syst.3
2022 Featured Cover
abstract
The cover image is based on the Research Article Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment by Zhiming Zhao et al., https://doi.org/10.1002/spe.3098.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.1
2022 Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment
abstract
Abstract Virtual research environments (VREs) provide user‐centric support in the lifecycle of research activities, for example, discovering and accessing research assets or composing and executing application workflows. A typical VRE is often implemented as an integrated environment, including a catalog of research assets, a workflow management system, a data management framework, and tools for enabling user collaboration. In contrast, notebook environments like Jupyter allow researchers to rapidly prototype scientific code and share their experiments as online accessible notebooks. Jupyter can support several popular languages used by data scientists, such as Python, R, and Julia. However, such notebook environments do not have seamless support for running heavy computations on remote infrastructure or finding and accessing collaborative software code inside notebooks. This article investigates the gap between a notebook environment and a VRE and proposes an embedded VRE solution for the Jupyter environment called Notebook‐as‐a‐VRE (NaaVRE). The NaaVRE solution provides functional components via a component marketplace and allows users to create a customized VRE on top of the Jupyter environment. From the VRE, a user can search research assets (data, software, and algorithms), compose workflows, manage the lifecycle of an experiment, and share the results among users in the community. We demonstrate how such a solution can enhance a legacy workflow that uses Light Detection and Ranging (LiDAR) data from country‐wide airborne laser scanning surveys for deriving geospatial data products of ecosystem structure at high resolution over broad spatial extents. This enables users to scale out the processing of multi‐terabyte LiDAR point clouds for ecological applications to more data sources in a distributed cloud environment. Similar applications could be developed for workflows producing other essential biodiversity variables.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.1
2021 Towards A Robust Meta-Reinforcement Learning-Based Scheduling Framework for Time Critical Tasks in Cloud Environments
abstract
Container clusters play an increasingly important role in cloud computing for processing dynamic computing tasks. The resource manager (i.e., orchestrater) of the cluster automates the scheduling of the dynamic requests, effectively manages the resources' utilization across distributing infrastructure resources. For many applications, the requests to the cluster are often with restricted deadlines. The scheduling of container clusters is often tricky, especially when the cluster's size is large and the load of the requests is dynamically changing. Machine learning-based approaches such as reinforcement learning have attracted lots of research attention during the past years; However, those approaches suffer from low robustness when the requests in an operational environment are changing and different from the training data sets. This paper investigates this problem by quantifying the robustness and proposing meta-gradient reinforcement learning to improve the robustness of classical reinforcement learning-based approaches. The proposed approach can lead to better deadline guarantees and faster adaptation for time-critical task scheduling under dynamic environments. We then empirically test the benefits of our method using both real-world and synthetic data sets. The evaluation results show that the proposed method outperforms the compared RL methods in scheduling performance and robustness.
Hongyun Liu, Peng Chen 0007, Zhiming Zhao
CLOUD3
2021 Unsupervised Anomaly Detection in Data Quality Control
abstract
Data is one of the most valuable assets of an organization and has a tremendous impact on its long-term success and decision-making processes. Typically, organizational data error and outlier detection processes perform manually and reactively, making them time-consuming and prone to human errors. Additionally, rich data types, unlabeled data, and increased volume have made such data more complex. Accordingly, an automated anomaly detection approach is required to improve data management and quality control processes. This study introduces an unsupervised anomaly detection approach based on models comparison, consensus learning, and a combination of rules of thumb with iterative hyper-parameter tuning to increase data quality. Furthermore, a domain expert is considered a human in the loop to evaluate and check the data quality and to judge the output of the unsupervised model. An experiment has been conducted to assess the proposed approach in the context of a case study. The experiment results confirm that the proposed approach can improve the quality of organizational data and facilitate anomaly detection processes.
Lex Poon, Siamak Farshidi, Zhiming Zhao
IEEE BigData4
2021 Blockchain-based prosumer incentivization for peak mitigation through temporal aggregation and contextual clustering
abstract
Peak mitigation is of interest to power companies as peak periods may require the operator to over provision supply in order to meet the peak demand. Flattening the usage curve can result in cost savings, both for the power companies and the end users. Integration of renewable energy into the energy infrastructure presents an opportunity to use excess renewable generation to supplement supply and alleviate peaks. In addition, demand side management can shift the usage from peak to off-peak times and reduce the magnitude of peaks. In this work, we present a data driven approach for incentive-based peak mitigation. Understanding user energy profiles is an essential step in this process. We begin by analysing a popular energy research dataset published by the Ausgrid corporation. Extracting aggregated user energy behavior in temporal contexts and semantic linking and contextual clustering give us insight into consumption and rooftop solar generation patterns. We implement, and performance test a blockchain-based prosumer incentivization system. The smart contract logic is based on our analysis of the Ausgrid dataset. Our implementation is capable of supporting 792,540 customers with a reasonably low infrastructure footprint.
Nikita Karandikar, Rockey Abhishek, Nishant Saurabh, Zhiming Zhao, Alexander Lercher, Ninoslav Marina, Radu Prodan, Chunming Rong, Antorweep Chakravorty
Blockchain Res. Appl.4
2021 The ARTICONF approach to decentralized car-sharing
abstract
Social media applications are essential for next-generation connectivity. Today, social media are centralized platforms with a single proprietary organization controlling the network and posing critical trust and governance issues over the created and propagated content. The ARTICONF project funded by the European Union's Horizon 2020 program researches a decentralized social media platform based on a novel set of trustworthy, resilient and globally sustainable tools that address privacy, robustness and autonomy-related promises that proprietary social media platforms have failed to deliver so far. This paper presents the ARTICONF approach to a car-sharing decentralized application (DApp) use case, as a new collaborative peer-to-peer model providing an alternative solution to private car ownership. We describe a prototype implementation of the car-sharing social media DApp and illustrate through real snapshots how the different ARTICONF tools support it in a simulated scenario.
Nishant Saurabh, Carlos Rubia, Anandakumar Palanisamy, Spiros Koulouzis, Mirsat Sefidanoski, Antorweep Chakravorty, Zhiming Zhao, Aleksandar Karadimce, Radu Prodan
Blockchain Res. Appl.7
2021 Measuring success for a future vision: Defining impact in science gateways/virtual research environments
abstract
Summary Scholars worldwide leverage science gateways/virtual research environments (VREs) for a wide variety of research and education endeavors spanning diverse scientific fields. Evaluating the value of a given science gateway/VRE to its constituent community is critical in obtaining the financial and human resources necessary to sustain operations and increase adoption in the user community. In this article, we feature a variety of exemplar science gateways/VREs and detail how they define impact in terms of, for example, their purpose, operation principles, and size of user base. Further, the exemplars recognize that their science gateways/VREs will continuously evolve with technological advancements and standards in cloud computing platforms, web service architectures, data management tools and cybersecurity. Correspondingly, we present a number of technology advances that could be incorporated in next‐generation science gateways/VREs to enhance their scope and scale of their operations for greater success/impact. The exemplars are selected from owners of science gateways in the Science Gateways Community Institute (SGCI) clientele in the United States, and from the owners of VREs in the International Virtual Research Environment Interest Group (VRE‐IG) of the Research Data Alliance. Thus, community‐driven best practices and technology advances are compiled from diverse expert groups with an international perspective to envisage futuristic science gateway/VRE innovations.
Prasad Calyam, Nancy Wilkins-Diehr, Mark A. Miller, Emre H. Brookes, Ritu Arora, Amit Chourasia, Douglas M. Jennewein, Viswanath Nandigam, Michael Drew Lamar, Sean B. Cleveland, Greg Newman, Shaowen Wang 0001, Ilya Zaslavsky, Michael A. Cianfrocco, Kevin M. Ellett, David G. Tarboton, Keith G. Jeffery, Zhiming Zhao, Juan González-Aranda, Mark J. Perri, Gregory E. Tucker, Leonardo Candela, Tamás Kiss, Sandra Gesing
Concurr. Comput. Pract. Exp.18
2021 Distributed service-level agreement management with smart contracts and blockchain
abstract
Summary The current cloud market is dominated by a few providers, which offer cloud services in a take‐it‐or‐leave‐it manner. However, the dynamism and uncertainty of cloud environments may require the change over time of both application requirements and service capabilities. The current service‐level agreement (SLA) management solutions cannot easily guarantee a trustworthy, distributed SLA adaptation due to the centralized authority of the cloud provider who could also misbehave to pursue individual goals. To address the above issues, we propose a novel SLA management framework, which facilitates the specification and enforcement of dynamic SLAs that enable one to describe how, and under which conditions, the offered service level can change over time. The proposed framework relies on a two‐level blockchain architecture. At the first level, the smart SLA is transformed into a smart contract that dynamically guides service provisioning. At the second level, a permissioned blockchain is built through a federation of monitoring entities to generate objective measurements for the smart SLA/contract assessment. The scalability of this permissioned blockchain is also thoroughly evaluated. The proposed framework enables creating open distributed clouds, which offer manageable and dynamic services, and facilitates cost reduction for cloud consumers, while it increases flexibility in resource management and trust in the offered cloud services.
Rafael Brundo Uriarte, Huan Zhou 0006, Kyriakos Kritikos, Zeshun Shi, Zhiming Zhao, Rocco De Nicola
Concurr. Comput. Pract. Exp.5
2021 Enforcing trustworthy cloud SLA with witnesses: A game theory-based model using smart contracts
abstract
There lacks trust between the cloud customer and provider to enforce traditional cloud SLA (Service Level Agreement) where the blockchain technique seems a promising solution. However, current explorations still face challenges to prove that the off-chain SLO (Service Level Objective) violations really happen before recorded into the on-chain transactions. In this paper, a witness model is proposed implemented with smart contracts to solve this trust issue. The introduced role, "Witness", gains rewards as an incentive for performing the SLO violation report, and the payoff function is carefully designed in a way that the witness has to tell the truth, for maximizing the rewards. This fact that the witness has to be honest is analyzed and proved using the Nash Equilibrium principle of game theory. For ensuring the chosen witnesses are random and independent, an unbiased selection algorithm is proposed to avoid possible collusions. An auditing mechanism is also introduced to detect potential malicious witnesses. Specifically, we define three types of malicious behaviors and propose quantitative indicators to audit and detect these behaviors. Moreover, experimental studies based on Ethereum blockchain demonstrate the proposed model is feasible, and indicate that the performance, ie, transaction fee, of each interface follows the design expectations.
Huan Zhou 0006, Xue Ouyang 0003, Jinshu Su, Cees T. A. M. de Laat, Zhiming Zhao
Concurr. Comput. Pract. Exp.5
2021 A Cost-Quality Beneficial Cell Selection Approach for Sparse Mobile Crowdsensing With Diverse Sensing Costs
abstract
The Internet of Things (IoT) and mobile techniques enable real-time sensing for urban computing systems. By recruiting only a small number of users to sense data from selected subareas (namely, cells), sparse mobile crowdsensing (MCS) emerges as an effective paradigm to reduce sensing costs for monitoring the overall status of a large-scale area. The current sparse MCS solutions reduce the sensing subareas (by selecting the most informative cells) based on the assumption that each sample has the same cost, which is not always realistic in the real world, as the cost of sensing in a subarea can be diverse due to many factors, e.g., the condition of the device, location, and routing distance. To address this issue, we proposed a new cell selection approach consisting of three steps (information modeling, cost estimation, and cost-quality beneficial cell selection) to further reduce the total costs and improve the task quality. Specifically, we discussed the properties of the optimization goals and modeled the cell selection problem as a solvable biobjective optimization problem under certain assumptions and approximations. Then, we presented two selection strategies, i.e., the Pareto optimization selection (POS) and generalized cost-benefit greedy (GCB-GREEDY) selection along with our proposed cell selection algorithm. Finally, the superiority of our cell selection approach is assessed through four real-life urban monitoring data sets (Parking, Flow, Traffic, and Humidity) and three cost maps (independent identically distributed with dynamic cost map, monotonic with dynamic cost map, and spatial-correlated cost map). Results show that our proposed selection strategies POS and GCB-GREEDY can save up to 15.2% and 15.02% sample costs and reduce the inference errors to a maximum of 16.8% (15.5%) compared to the baseline-query by committee (QBC) in a sensing cycle. The findings show important implications in sparse MCS for urban context properties.
Zhengqiu Zhu, Bin Chen 0003, Zhong Liu 0002, Zhiming Zhao
IEEE Internet Things J.6
2021 Building a blockchain-based decentralized ecosystem for cloud and edge computing: an ALLSTAR approach and empirical study
Huan Zhou 0006, Zeshun Shi, Xue Ouyang 0003, Zhiming Zhao
Peer-to-Peer Netw. Appl.4
2020 Sharing digital object across data infrastructures using Named Data Networking (NDN)
abstract
Data infrastructures manage the life cycle of digital assets and allow users to efficiently discover them. To improve the Findability, Accessibility, Interoperability and Re-usability (FAIRness) of digital assets, a data infrastructure needs to provide digital assets with not only rich meta information and semantics contexts information but also globally resolvable identifiers. The Persistent Identifiers (PIDs), like Digital Object Identifier (DOI) are often used by data publishers and infrastructures. The traditional IP network and client-server model can potentially cause congestion and delays when many consumers simultaneously access data. In contrast, Information Centric Networking (ICN) technologies such as Named Data Networking (NDN) adopt a data centric approach where digital data objects, once requested, may be stored on intermediate hops in the network. Consecutive requests for that unique digital object are then made available by these intermediate hops (caching). This approach distributes traffic load more efficient and reliable compared to host-to-host connection oriented techniques, and demonstrates attractive opportunities for sharing digital objects across distributed networks. However, such an approach also faces several challenges. It requires not only an effective translation between the different naming schemas among PIDs and NDN, in particular for supporting PIDs from different publishers or repositories. Moreover, the planning and configuration of an ICN environment for distributed infrastructures are lacking an automated solution. To bridge the gap, we propose an ICN planning service with specific consideration of interoperability across PID schemas in the Cloud environment.
Kees de Jong, Cas Fahrenfort, Anas Younis, Zhiming Zhao
CCGRID4
2020 A Trustworthy Blockchain-based Decentralised Resource Management System in the Cloud
abstract
Quality Critical Decentralised Applications (QC-DApp) have high requirements for system performance and service quality, involve heterogeneous infrastructures (Clouds, Fogs, Edges and IoT), and rely on the trustworthy collaborations among participants of data sources and infrastructure providers to deliver their business value. The development of the QCDApp has to tackle the low-performance challenge of the current blockchain technologies due to the low collaboration efficiency among distributed peers for consensus. On the other hand, the resilience of the Cloud has enabled significant advances in software-defined storage, networking, infrastructure, and every technology; however, those rich programmabilities of infrastructure (in particular, the advances of new hardware accelerators in the infrastructures) can still not be effectively utilised for QCDApp due to lack of suitable architecture and programming model.
Zhiming Zhao, Chunming Rong, Martin Gilje Jaatun
ICPADS1
2020 ALLSTAR: A Blockchain Based Decentralized Ecosystem for Cloud and Edge Computing
abstract
Last decades, Cloud computing has made significant impacts on traditional applications to change their development and operation methods. We witnessed ever more newly-built Clouds and data centers. However, the centralized management mechanism of current Clouds lacks the dispersion to satisfy the requirements of emerging collaborative applications, including AI, IoT, and autopilot. On the other hand, the Edge computing stays at the conceptual and experimental stage. Most organizations construct their own Edge nodes to operate applications. An efficient and incentive mechanism is missing to motivate the Edge and micro Cloud resource providers to join and constitute a more generalized and decentralized ecosystem. To address this issue, we propose ALLSTAR, a blockchain based architecture for equally combining all the Cloud and Edge resources to be seamlessly leveraged by the application in the DevOps (development and operations) lifecycle. The ALLSTAR architecture is a systematic solution to realize the "Cloud+Edge" management and contributes to constructing the corresponding ALLSTAR ecosystem. This paper describes the overall architecture of ALLSTAR, the related key techniques, and detailed application DevOps processes as well as the new business model.
Huan Zhou 0006, Xue Ouyang 0003, Zhiming Zhao
JCC3
2020 Decentralized workflow management on software defined infrastructures
abstract
Data-intensive workflow applications are characterized by their continuously growing volumes of data being processing, the complexity of tasks in the pipeline, and infrastructure capacity required for computation and storage. The infrastructure technologies of computing, storage and networking have made tremendous progress during the past yeas. We review the emerging trends in the data-intensive workflow applications, in particular the potential challenges and opportunities enabled by the decentralized application paradigm.
Yuandou Wang, Zhiming Zhao
SERVICES2
2020 Time-critical data management in clouds: Challenges and a Dynamic Real-Time Infrastructure Planner (DRIP) solution
abstract
Summary The increasing volume of data being produced, curated, and made available by research infrastructures in the environmental science domain require services that are able to optimize the delivery and staging of data for researchers and other users of scientific data. Specialized data services for managing data life cycle, for creating and delivering data products, and for customized data processing and analysis all play a crucial role in how these research infrastructures serve their communities, and many of these activities are time‐critical—needing to be carried out frequently within specific time windows. We describe our experiences identifying the time‐critical requirements of environmental scientists making use of computational research support environments. We present a microservice‐based infrastructure optimization suite, the Dynamic Real‐Time Infrastructure Planner, used for constructing virtual infrastructures for research applications on demand. We provide a case study whereby our suite is used to optimize runtime service quality for a data subscription service provided by the Euro‐Argo using EGI Federated Cloud and EUDAT's B2SAFE services, and to consider how such a case study relates to other application scenarios.
Spiros Koulouzis, Paul Martin 0002, Huan Zhou 0006, Yang Hu 0013, Thierry Carval, Baptiste Grenier, Jani Heikkinen, Cees T. A. M. de Laat, Zhiming Zhao
Concurr. Comput. Pract. Exp.10
2020 Concurrent container scheduling on heterogeneous clusters with multi-resource constraints
abstract
By effectively virtualizing operating systems and encapsulating necessary runtime contexts of software components and services, container technologies can significantly improve portability and efficiency for distributed application deployment. It flexibly extends virtual machine based cloud (Infrastructure-as-a-Service) as a much lighter virtual environment (container cluster) for agile application management. However, existing container management systems are not capable of handling concurrent requests efficiently, particularly for the underlying clusters with heterogeneous machines and the requested containers with multi-resource demands. In this paper, we propose an Enhanced Container Scheduler (ECSched) for efficiently scheduling concurrent container requests on heterogeneous clusters with multi-resource constraints. We formulate the container scheduling problem as a minimum cost flow problem (MCFP), and represent the container requirements using a specific graph data structure (flow network). ECSched affords flexibility in constructing the flow network based on a batch of concurrent requests, and performs the MCFP algorithm to schedule the concurrent requests in an online manner. We evaluate ECSched in different testbed clusters, and measure the scheduling overhead with large-scale simulations. The experimental results show that ECSched outperforms state-of-the-art container schedulers in container performance and resource efficiency, and only introduces a small and acceptable scheduling overhead in large-scale clusters.
Yang Hu 0013, Huan Zhou 0006, Cees T. A. M. de Laat, Zhiming Zhao
Future Gener. Comput. Syst.4
2020 Editorial for FGCS Special issue on "Time-critical Applications on Software-defined Infrastructures"
Zhiming Zhao, Ian J. Taylor, Radu Prodan
Future Gener. Comput. Syst.1
2019 Multi-objective Container Deployment on Heterogeneous Clusters
abstract
Operating system (OS) containers are becoming increasingly popular in cloud computing for improving productivity and code portability. However, existing deployment scheduling solutions mainly treat each container deployment as an independent request, and focus on the single aspect of resource utilization or load balancing, or work on homogeneous clusters. In this paper, we propose a new container deployment algorithm to satisfy multiple objectives on heterogeneous clusters. We analyze the deployment requirements of container-based infrastructure and formulate the deployment problem as a vector bin packing problem with heterogeneous bins. We focus on three objectives: multi-resource guarantee, load balancing, and dependency awareness. The goal of the proposed algorithm is to improve the tradeoff between load balancing and dependency awareness with multi-resource guarantees. Based on the algorithm, we implement a prototype scheduler to deploy containers on heterogeneous clusters. We evaluate our scheduler over a wide range of workload scenarios by simulation, which shows that our scheduler significantly outperforms existing schedulers of the container orchestration platforms.
Yang Hu 0013, Cees T. A. M. de Laat, Zhiming Zhao
CCGRID3
2019 An Automated Customization and Performance Profiling Framework for Permissioned Blockchains in a Virtualized Environment
abstract
The permissioned blockchains have demonstrated their potential to provide trustworthy and security services in various industrial scenarios, especially in the Cloud-based virtualized environments. To customize the configuration of a blockchain application, an operator needs the performance characteristics of a blockchain network in different Cloud environments. However, manually profiling the performance characteristics of a blockchain network is very time-consuming. Therefore, in this paper, we propose a BlockchaIn-infRAstructure CustomIzation and Auto-profiLing (BIRACIAL) framework to automate the whole process of blockchain deployment and performance profiling. Based on the profile and performance requirements of a blockchain application, the framework aims to plan the virtual infrastructure for permissioned blockchain, to automate the provision of the required infrastructure, to deploy the customized permissioned blockchain, and to enable continuous monitoring of blockchain performance. Our evaluation results show that the proposed framework can achieve automated deployment of different permissioned blockchain networks under certain overheads. The performance profiling results can be used to compare and select the appropriate blockchain platforms and consensus algorithms.
Zeshun Shi, Huan Zhou 0006, Jayachander Surbiryala, Yang Hu 0013, Cees T. A. M. de Laat, Zhiming Zhao
CloudCom6
2019 Contextual Linking between Workflow Provenance and System Performance Logs
abstract
When executing scientific workflows, anomalies of the workflow behavior are often caused by different issues such as resource failures at the underlying infrastructure. The provenance information collected by workflow management systems only captures the transformation of data at the workflow level. Analyzing provenance information and apposite system metrics requires expertise and manual effort. Moreover, it is often time-consuming to aggregate this information and correlate events occurring at different levels of the infrastructure. In this paper, we propose an architecture to automate the integration among workflow provenance information and performance information from the infrastructure level. Our architecture enables workflow developers or domain scientists to effectively browse workflow execution information together with the system metrics, and analyze contextual information for possible anomalies.
Elias el Khaldi Ahanach, Spiros Koulouzis, Zhiming Zhao
eScience3
2019 Teaching DevOps and Cloud Based Software Engineering in University Curricula
abstract
This paper presents recommendations on the design and pilot implementation of the DevOps and Cloud based Software Development curricula for Computer Science and Software Engineering masters. The central part of proposed approach is the Body of Knowledge in the DevOps technologies for Software Engineering (DevOpsSE BoK) that defines a set Knowledge Areas and Knowledge Units required for SE professionals to work efficiently as DevOps engineer or application developer. Defining DevOpsSE-BoK provides a basis for defining required professional competences and skills and allows consistent curricula structuring and profiling. The paper also reports on the experience of the first course run on 2018/2019 academic year at the University of Amsterdam. The paper presents the structure of the course and explains what instructional methodologies have been used for course development, such as project based learning that facilitates the students' team based skills both in mastering Agile development process and skills sharing. The paper provides a short summary of the generally used DevOps definitions, concepts, models and tools, specifically focusing on the cloud based DevOps tools for software development, deployment and operation that allows the main DevOps principle of continuous development and continuous improvement which are critical for modern agile data driven companies.
Yuri Demchenko, Zhiming Zhao, Jayachander Surbiryala, Spiros Koulouzis, Zeshun Shi, Jelena Gordiyenko
eScience2
2019 Effective Digital Object Access and Sharing Over a Networked Environment using DOIP and NDN
abstract
FAIRness (findability, accessibility, interoperability and re-usability) is crucial for enabling open science and innovation based on digital objects from large communities of providers and users. However, the gaps among version control, identification and distributed access systems often make the scalability of data centric applications difficult across large user communities and highly distributed infrastructures. This poster proposes a solution for accessing and sharing digital objects over a networked environment using Digital object interface protocol (DOIP) and Named Data Networking (NDN).
Cas Fahrenfort, Zhiming Zhao
eScience2
2019 ENVRI-FAIR - Interoperable Environmental FAIR Data and Services for Society, Innovation and Research
abstract
ENVRI-FAIR is a recently launched project of the European Union's Horizon 2020 program (EU H2020), connecting the cluster of European Environmental Research Infrastructures (ENVRI) to the European Open Science Cloud (EOSC). The overarching goal of ENVRI-FAIR is that all participating research infrastructures (RIs) will provide a set of interoperable FAIR data services that enhance the efficiency and productivity of researchers, support innovation, enable data-and knowledge-based decisions and connect the ENVRI cluster to the EOSC. This goal will be reached by: (1) defining community policies and standards across all stages of the data life cycle, aligned with the wider European policies and with international developments; (2) creating for all participating RIs sustainable, transparent and auditable data services for each stage of the data life cycle, following the FAIR principles; (3) implementing prototypes for testing pre-production services at each RI, leading to a catalogue of prepared services; (4) exposing the complete set of thematic data services and tools of the ENVRI cluster to the EOSC catalogue of services.
Andreas Petzold, Ari Asmi, Alex Vermeulen, Gelsomina Pappalardo, Daniele Bailo, Dick Schaap, Helen M. Glaves, Ulrich Bundke, Zhiming Zhao
eScience9
2019 A Blockchain based Witness Model for Trustworthy Cloud Service Level Agreement Enforcement
abstract
Traditional cloud Service Level Agreement (SLA) suffers from lacking a trustworthy platform for automatic enforcement. The emerging blockchain technique brings in an immutable solution for tracking transactions among business partners. However, it is still very challenging to prove the credibility of possible violations in the SLA before recording them onto the blockchain. To tackle this challenge, we propose a witness model using game theory and the smart contract techniques. The proposed model extends the existing service model with a new role called “witness” for detecting and reporting service violations. Witnesses gain revenue as an incentive for performing these duties, and the payoff function is carefully designed in a way that trustworthiness is guaranteed: in order to get the maximum profit, the witness has to always tell the truth. This is analyzed and proved through game theory using the Nash equilibrium principle. In addition, an unbiased sortition algorithm is proposed to ensure the randomness of the independent witnesses selection from the decentralized witness pool, to avoid possible unfairness or collusion. An auditing mechanism is also introduced in the paper to detect potential irrational or malicious witnesses. We have prototyped the system leveraging the smart contracts of Ethereum blockchain. Experimental results demonstrate the feasibility of the proposed model and indicate good performance in accordance with the design expectations.
Huan Zhou 0006, Xue Ouyang 0003, Zhijie Ren, Jinshu Su, Cees T. A. M. de Laat, Zhiming Zhao
INFOCOM6
2019 Operating Permissioned Blockchain in Clouds: A Performance Study of Hyperledger Sawtooth
abstract
With ever more IoT (Internet of Things) and bigdata applications, the emerging blockchain techniques provide fundamental supports to credibly track the transactions of digital assets. Public blockchains, e.g., bitcoin, are often energy-consuming and low efficient. Therefore, an empirical study of operating permissioned blockchains in clouds is urgently needed. In this paper, we study the performance of Sawtooth, a well-known permissioned blockchain platforms from Hyperledger, in cloud environments. Our results provide insights for blockchain operators to optimize the performance of Sawtooth through adjusting the two configuration parameters, i.e., Scheduler and Maximum Batches Per Block. Our approach can be used to test other blockchain platforms.
Zeshun Shi, Huan Zhou 0006, Yang Hu 0013, Jayachander Surbiryala, Cees T. A. M. de Laat, Zhiming Zhao
ISPDC6
2019 Learning Workflow Scheduling on Multi-Resource Clusters
abstract
Workflow scheduling is one of the key issues in the management of workflow execution. Typically, a workflow application can be modeled as a Directed-Acyclic Graph (DAG). In this paper, we present GoDAG, an approach that can learn to well schedule workflows on multi-resource clusters. GoDAG directly learns the scheduling policy from experience through deep reinforcement learning. In order to adapt deep reinforcement learning methods, we propose a novel state representation, a practical action space and a corresponding reward definition for workflow scheduling problem. We implement a GoDAG prototype and a simulator to simulate task running on multi-resource clusters. In the evaluation, we compare the GoDAG with three state-of-the-art heuristics. The results show that GoDAG outperforms the baseline heuristics, leading to less average makespan to different workflow structures.
Yang Hu 0013, Cees T. A. M. de Laat, Zhiming Zhao
NAS3
2019 Knowledge-as-a-Service: A Community Knowledge Base for Research Infrastructures in Environmental and Earth Sciences
abstract
The ENVRI Reference Model (ENVRI RM) and its ontological representation Open Information Linking for Environmental RIs (OIL-E) allow architects and engineers to describe the architecture and operational behavior of environmental and earth science research infrastructures (RIs) in a standardized way using community-agreed terminology. RI descriptions can be published as linked data, allowing discovery, querying and comparison using established Semantic Web technologies. The ENVRI Knowledge Base is a community knowledge base which uses OIL-E to capture information about environmental and earth science RIs in the ENVRI community for query and comparison. Such Knowledge-as-a-Service supports identifying the technologies and standards used for particular activities and services and evaluating research infrastructure subsystems and behaviors against certain criteria, such as compliance with the FAIR data principles.
Zhiming Zhao, Paul Martin 0002, Jordan Maduro, Peter Thijsse, Dick Schaap, Markus Stocker, Doron Goldfarb, Barbara Magagna
SERVICES1
2019 Mapping heterogeneous research infrastructure metadata into a unified catalogue for use in a generic virtual research environment
Paul Martin 0002, Laurent Remy, Maria Theodoridou, Keith G. Jeffery, Zhiming Zhao
Future Gener. Comput. Syst.5
2019 SWITCH workbench: A novel approach for the development and deployment of time-critical microservice-based cloud-native applications
Polona Stefanic, Matej Cigale, Andrew C. Jones, Louise Knight, Ian J. Taylor, Cristiana Istrate, George Suciu, Alexandre Ulisses, Vlado Stankovski, Salman Taherizadeh, Guadalupe Flores Salado, Spiros Koulouzis, Paul Martin 0002, Zhiming Zhao
Future Gener. Comput. Syst.14
2019 Profiling the scheduling decisions for handling critical paths in deadline-constrained cloud workflows
abstract
In this paper, we study the scheduling decisions for handling deadline-constrained workflows in the context of planning customized virtual infrastructures in the cloud. We specifically focus on the effects of using different types of greediness in selecting cost-effective virtual machines for the tasks in an application’s workflow graph. The profiling procedure followed demonstrates that for the widely used approach of the partial critical path algorithm a greedy version is preferred to a more stringent version under different stress conditions, from tight to loose deadlines. Representative topologies of workflow applications are used to generate sets of task graph scheduling problems. Monitoring the performance of the partial critical path algorithm with different types of greediness reveals which of the topologies tested are difficult to solve under various stress conditions. It turns out that an invalid outcome of a greedy version of the partial critical path algorithm is more susceptible to become valid via a final refinement cycle than a less greedy version. The procedure outlined in this paper will allow for a systematic study of a specific heuristic in a workflow scheduling method to increase its success in infrastructure planning under different deadline conditions and is proposed to be part of a general profiling framework.
Arie Taal, Cees T. A. M. de Laat, Zhiming Zhao
Future Gener. Comput. Syst.4
2019 CloudsStorm: A framework for seamlessly programming and controlling virtual infrastructure functions during the DevOps lifecycle of cloud applications
abstract
Summary The infrastructure‐as‐a‐service (IaaS) model of cloud computing provides virtual infrastructure functions (VIFs), which allow application developers to flexibly provision suitable virtual machines' (VM) types and locations, and even configure the network connection for each VM. Because of the pay‐as‐you‐go business model, IaaS provides an elastic way to operate applications on demand. However, in current cloud applications DevOps (software development and operations) lifecycle, the VM provisioning steps mainly rely on manually leveraging these VIFs. Moreover, these functions cannot be programmatically embedded into the application logic to control the infrastructure at runtime. Especially, the vendor lock‐in issue, which different clouds provide different VIFs, also enlarges this gap between the cloud infrastructure management and application operation. To mitigate this gap, we designed and implemented a framework, CloudsStorm, which enables developers to easily leverage VIFs of different clouds and program them into their cloud applications. To be specific, CloudsStorm empowers applications with infrastructure programmability at design‐level, infrastructure‐level, and application‐level. CloudsStorm also provides two infrastructure controlling modes, ie, active and passive mode, for applications at runtime. Besides, case studies about operating task‐based and big data applications on clouds show that the monetary cost is significantly reduced through the seamless and on‐demand infrastructure management provided by CloudsStorm. Finally, the scaling and recovery operation evaluations of CloudsStorm are performed to show its controlling performance. Compared with other tools, ie, “jcloud” and “cloudinit.d”, the scaling and provisioning performance evaluations demonstrate that CloudsStorm can achieve at least 10% efficiency improvement in our experiment settings.
Huan Zhou 0006, Yang Hu 0013, Xue Ouyang 0003, Jinshu Su, Spiros Koulouzis, Cees T. A. M. de Laat, Zhiming Zhao
Softw. Pract. Exp.7
2018 Empowering Dynamic Task-Based Applications with Agile Virtual Infrastructure Programmability
abstract
The IaaS (Infrastructure-as-a-Service) offered by Clouds provides applications with the capability of customizing VMs and configuring their network. Compared to traditional service-based IaaS applications such as persistent web services, most task-based applications have a relatively short duration but are triggered on demand. A typical way to support such kinds of application is to provision a shared and fixed virtual infrastructure based on pre-estimated size in advance, and then perform all the processing tasks. However, due to unpredictable workloads, this solution can lead to either cost inefficiency caused by over-provisioning, or failure to deliver the performance required by applications. CloudsStorm is a dynamic control framework proposed to provide applications with agile programmability and flexibility in controlling the virtual infrastructure. With its front end, applications can design their networked infrastructure and program that infrastructure with our interpreted infrastructure code language. With the back-end engine, the infrastructure code can be executed to provision the networked infrastructure, deploy and execute the application to obtain results, and release resources. Moreover, we adopt multi-threading to support parallel operation. Finally, we conduct experiments in an assumed scenario to demonstrate functionalities of CloudsStorm. The evaluation results prove CloudsStorm is efficient for task-based applications that need to exploit Clouds but reduce the monetary cost.
Huan Zhou 0006, Yang Hu 0013, Jinshu Su, Mingmin Chi, Cees T. A. M. de Laat, Zhiming Zhao
IEEE CLOUD6
2018 Information Centric Networking for Sharing and Accessing Digital Objects with Persistent Identifiers on Data Infrastructures
abstract
Persistent identifiers (PIDs) such as Digital Object Identifiers (DOIs) provide a unique and persistent way to identify and cite digital objects such as publications, media content and research data. They are widely used by data producers to catalogue and publish digital assets and research data. Nowadays, research infrastructures (RIs) offer services not only for accessing and publishing data objects, but also for processing data based on user demands, e.g., via scientific workflows or third party virtual research environments. However, efficiently retrieving and sharing digital objects in a shared data processing environment requires knowledge of application access patterns as well as the underlying network level distribution. As the number and size of data objects increases, optimizing data discovery and access among distributed partners on shared infrastructure emerges as an important challenge for infrastructure operators to maintain quality of service and user experience. In this paper, we propose a novel approach that utilizes Information Centric Networking (ICN) to retrieve content based on PIDs while optimizing data access on shared infrastructure.
Spiros Koulouzis, Rahaf Mousa, Andreas Karakannas, Cees T. A. M. de Laat, Zhiming Zhao
CCGrid5
2018 Trustworthy Cloud Service Level Agreement Enforcement with Blockchain Based Smart Contract
abstract
Cloud Service Level Agreement (SLA) is challengeable due to lacking a trustworthy platform. This paper presents a witness model to credibly enforce the cloud service level agreement. Through introducing the witness role and using the blockchain based smart contract, we solve the trust issues about who can detect the service violation, how the violation is confirmed and the compensation is guaranteed. In this model, a verifiable consensus sortition algorithm proposed by us is firstly leveraged to select independent witnesses to form a witness committee. They are responsible for a specific service level agreement and get paid by monitoring and detecting service violation. Through carefully designing the witness' payoff function in the agreement, we further leverage game theory to analyze and prove that it is not the witness itself is trustworthy. Instead, the witness has to tell the truth because of its greedy nature, which is the desire to maximize its own revenue. As long as the service violation is confirmed by the witness committee, the compensation is automatically transferred to the customer by the smart contract. Finally, we implement a proof-of-concept prototype with the smart contract of Ethereum blockchain. It demonstrates the feasibility of our model.
Huan Zhou 0006, Cees T. A. M. de Laat, Zhiming Zhao
CloudCom3
2018 ECSched: Efficient Container Scheduling on Heterogeneous Clusters
Yang Hu 0013, Huan Zhou 0006, Cees T. A. M. de Laat, Zhiming Zhao
Euro-Par4
2018 Classification of High Resolution Urban Remote Sensing Images Using Deep Networks by Integration of Social Media Photos
abstract
In recent decades, it is easy to obtain remote sensing images which have been successfully applied to various applications, such as urban planning, hazard monitoring, etc. In particular, high resolution (HR) remote sensing (RS) images can better monitor our living environment from a broader spatial perspective. However, raw remote sensing images provide no labeling information to train a classifier, which usually is exploited to generate remote sensing maps. Based on our previous work, in the paper, an automatic classification system is proposed to classify high resolution urban RS images using deep neural networks, in particular, convolutional neural networks and fully convolutional networks. The labeling information is assigned on the context of both social media photos and HR remote sensing images by significantly reducing the cost of manual labeling without the necessity of remote sensing experts. The experiments carried out on high resolution remote sensing images acquired in the city Frankfurt taken by the Jilin-1 satellites confirm the effectiveness of the proposed strategy compared to the state of the art.
Yiqing Qin, Mingmin Chi, Yijian Zeng, Zhiming Zhao
IGARSS6
2018 A novel parallel distance metric-based approach for diversified ranking on large graphs
Jin Li 0007, Yun Yang 0003, Xiaoling Wang 0004, Zhiming Zhao, Tong Li 0004
Future Gener. Comput. Syst.4
2018 Monitoring self-adaptive applications within edge computing frameworks: A state-of-the-art review
abstract
Recently, a promising trend has evolved from previous centralized computation to decentralized edge computing in the proximity of end-users to provide cloud applications. To ensure the Quality of Service (QoS) of such applications and Quality of Experience (QoE) for the end-users, it is necessary to employ a comprehensive monitoring approach. Requirement analysis is a key software engineering task in the whole lifecycle of applications; however, the requirements for monitoring systems within edge computing scenarios are not yet fully established. The goal of the present survey study is therefore threefold: to identify the main challenges in the field of monitoring edge computing applications that are as yet not fully solved; to present a new taxonomy of monitoring requirements for adaptive applications orchestrated upon edge computing frameworks; and to discuss and compare the use of widely-used cloud monitoring technologies to assure the performance of these applications. Our analysis shows that none of existing widely-used cloud monitoring tools yet provides an integrated monitoring solution within edge computing frameworks. Moreover, some monitoring requirements have not been thoroughly met by any of them.
Salman Taherizadeh, Andrew C. Jones, Ian J. Taylor, Zhiming Zhao, Vlado Stankovski
J. Syst. Softw.4
2017 Deadline-Aware Coflow Scheduling in a DAG
abstract
Data-intensive applications usually need to deal with huge volumes of data within their deadlines. These applications can be modelled as DAGs and require parallel computation frameworks such as MapReduce and Spark to enhance the performance. The network communication has a crucial impact on the performance of an application. Coflow is intended to address the application-specific network level Quality-of-Service (QoS) requirements in cloud-based data centres. However, existing works mainly focus on scheduling coflows in a single stage. How to schedule coflows in multi-stage applications (represented as DAGs) remains to be an open problem. In this paper we study the problem of scheduling coflows in a DAG to meet its deadline requirement. Single stage coflow scheduling has been proven to be NP-hard. Multiple stages in a DAG make our problem even more complex. Owing to the complexity of the problem, we propose a genetic algorithm-based method for solving the problem. The effectiveness of our solution is verified through numerical evaluation. Experimental results show that our solution can effectively guarantee the deadline of the DAGs compared with existing single stage coflow scheduling algorithms.
Huan Zhou 0006, Yang Hu 0013, Cees T. A. M. de Laat, Zhiming Zhao
CloudCom5
2017 Deadline-Aware Deployment for Time Critical Applications in Clouds
Yang Hu 0013, Huan Zhou 0006, Paul Martin 0002, Arie Taal, Cees T. A. M. de Laat, Zhiming Zhao
Euro-Par7
2017 QoS-aware virtual SDN network planning
abstract
Software Defined Networking (SDN) technologies provide applications opportunities to manipulate underlying network flows and topologies via network controllers during runtime. In cloud environments, networked virtual machines can be enhanced by SDN by providing applications with controllable infrastructures to meet system-level quality requirements; however, customizing a suitable network topology with optimally placed controller(s) for given quality requirements and workload characteristics is often not an easy task. We call such problem virtual SDN network planning problem. In this paper, a Topology-Controller planner (TCPlanner) is proposed for customizing the network topology and placing the controllers. Experiments with different scales of network show that our approach can effectively plan virtual SDN networks to meet the various QoS requirements and reduce costs.
Cees T. A. M. de Laat, Zhiming Zhao
IM3
2017 Automatic Collector for Dynamic Cloud Performance Information
abstract
When deploying an application in the cloud, a developer often wants to know which of the wide variety of cloud resources is best to use. Most cloud providers only provide static information about different cloud resources which is often not enough because static information does not take into account the hardware and software that is being used or the policy that has been applied by the cloud provider. Therefore, dynamic benchmarking of cloud resources is needed to find out how a certain workload is going to behave on a certain instance. However, benchmarking various cloud resources is a time consuming process. Thus, using a tool which automatically benchmarks various cloud resources will be of great use. In this paper, we present the Cloud Performance Collector, a modular cloud benchmarking tool aimed to automatically benchmark a wide variety of applications. To demonstrate the benefit of the tool, we did three experiments with three synthetic benchmark applications and one real-world application using the ExoGENI testbed.
Olaf Elzinga, Spiros Koulouzis, Arie Taal, Yang Hu 0013, Huan Zhou 0006, Paul Martin 0002, Cees T. A. M. de Laat, Zhiming Zhao
NAS9
2017 Planning virtual infrastructures for time critical applications with multiple deadline constraints
Arie Taal, Paul Martin 0002, Yang Hu 0013, Huan Zhou 0006, Jianmin Pang, Cees T. A. M. de Laat, Zhiming Zhao
Future Gener. Comput. Syst.8
2016 Fast Resource Co-provisioning for Time Critical Applications Based on Networked Infrastructures
abstract
Resource provisioning is a key step in the deployment of applications onto clouds. When some datacenter is not accessible or some part of the infrastructure is crashed, the provisioning mechanism is therefore essential for these applications to recover quickly from sudden failures, especially for time critical applications. However, most current solutions focus on the cloud provider's hardware to achieve the fast provisioning of cloud resources. This paper proposes a co-provisioning mechanism to partition the customer's cloud resource requests while preserving their connectivity. This mechanism uses a brokering approach that is totally transparent to both the customer and the cloud provider, specifically considering the network topology. We carry out experiments on an NIaaS (networked infrastructure-as-a-service) platform, called ExoGENI. Experimental results and data analysis show that this mechanism is feasible and can dramatically improve the speed of resource provisioning.
Huan Zhou 0006, Yang Hu 0013, Jinshu Su, Paul Martin 0002, Cees T. A. M. de Laat, Zhiming Zhao
CLOUD7
2016 On the Next Generations of Infrastructure-as-a-Services
abstract
Following the wide adoption by industry of the cloud computing technologies, we can talk about a second generation of cloud services and products that are currently under design phase. However, it is not yet clear how the third generation of cloud products and services of the next decade will look like, especially at the delivery level of Infrastructure-as-a-Service. In order to answer at least partially to such a challenging question, we initiated a literature overview and two surveys involving the members of a cluster of European research and innovation actions. The results are interpreted in this paper and a set of topics of interest for the third generation are identified.
Dana Petcu, Maria Fazio, Radu Prodan, Zhiming Zhao, Massimiliano Rak
CLOSER (1)4
2016 Fast and Dynamic Resource Provisioning for Quality Critical Cloud Applications
abstract
As many quality critical applications are migrating to clouds, Quality of Service (QoS) and Quality of Experience (QoE) have become vital properties for cloud applications. Therefore, the provisioning mechanism, which aims to make the virtual infrastructure recover from sudden failures quickly or adapt dynamic properties of applications, is essential. However, most current provisioning mechanisms focus on the cloud provider and are developed for specific hardware. This paper proposes a mechanism to partition a customer's cloud resource requests efficiently across multiple domains or clouds, while ensuring that the partitions are still connected with each other. This mechanism exploits networked infrastructure to make dynamic cloud resource provisioning as fast as possible. It works using a broker-based model that is transparent both to the customer and to the cloud provider. It is easy for customers to use and does not force providers to make any changes to their services. Moreover, the dynamic property makes the provisioned infrastructure better able to recover from failures quickly. We implement the mechanism and carry out experiments on ExoGENI, a networked infrastructure-as-a-service (NIaaS) platform. Comprehensive experimental results and theoretical analysis demonstrate that the mechanism we propose is feasible and can dramatically improve the speed of resource provisioning.
Huan Zhou 0006, Yang Hu 0013, Paul Martin 0002, Cees T. A. M. de Laat, Zhiming Zhao
ISORC6
2016 SDN-aware federation of distributed data
Spiros Koulouzis, Adam Belloum, Marian Bubak, Zhiming Zhao, Miroslav Zivkovic, Cees T. A. M. de Laat
Future Gener. Comput. Syst.4
2015 A Software Workbench for Interactive, Time Critical and Highly Self-Adaptive Cloud Applications (SWITCH)
abstract
Time critical applications have very high requirements on network and computing services, in particular on well-tuned software architecture with sophisticated optimisation on data communication. Their development is often customised to dedicated infrastructure, and system performance is difficult to maintain when infrastructure changes. This fatal weakness in existing architecture and software tools causes very high development costs, and makes it difficult to fully utilise the virtualised, programmable and quality-on-demand services provided by networked Clouds to improve the system productivity. The Software Workbench for Interactive, Time Critical and Highly self-adaptive Cloud applications (SWITCH) is a newly funded project by EU H2020 to address this urgent industrial need, it aims at improving the existing development and execution model of time critical applications by introducing a novel conceptual model called application-infrastructure co-programming and control model, in which application QoS/QoE together with the programmability and controllability of Cloud environments can be all included in the complete lifecycle of applications.
Zhiming Zhao, Arie Taal, Andrew C. Jones, Ian J. Taylor, Vlado Stankovski, Ignacio Garcia Vega, Francisco Jesus Hidalgo, George Suciu, Alexandre Ulisses, Cees T. A. M. de Laat
CCGRID1
2015 Open Information Linking for Environmental Research Infrastructures
abstract
Environmental research infrastructures (RIs) support data-intensive research by integrating large-scale sensor/observer networks with dedicated data curation services and analytical tools. However the diversity of scientific disciplines coupled with the lack of an accepted methodology for constructing new RIs inevitably leads to incompatibilities between the data models, metadata standards and service descriptions used by different RIs, inhibiting their usefulness for interdisciplinary research. In the absence of a common global ontology of science and infrastructure, these inconsistencies may best be counteracted by selectively bridging the semantics of the various vocabularies, standards and models used by the RIs at present. Open Information Linking for Environmental RIs (OIL-E) was developed within the FP7 project ENVRI to provide a framework for semantic linking of knowledge resources used by different environmental RIs. Built around a multi-viewpoint reference model ENVRI-RM, OIL-E is intended to act as a central exchange for linking information fragments and identifying gaps in the conceptual models of RIs.
Paul Martin 0002, Paola Grosso, Barbara Magagna, Herbert Schentz, Yin Chen 0004, Alex R. Hardisty, Wouter Los, Keith G. Jeffery, Cees T. A. M. de Laat, Zhiming Zhao
e-Science10
2015 Reference Model Guided System Design and Implementation for Interoperable Environmental Research Infrastructures
abstract
Environmental research infrastructures (RIs) support their respective research communities by integrating large-scale sensor/observation networks with data curation services, analytical tools and common operational policies. These RIs are developed as pillars of intra-and interdisciplinary research, however comprehension of the complex, pathologically interconnected aspects of the Earth's ecosystem increasingly requires that researchers conduct their experiments across infrastructure boundaries. Consequently, almost all data-related activities within these infrastructures, from data capture to data usage, needs to be designed to be broadly interoperable in order to enable real interdisciplinary innovation. The Data for Science theme in the EU Horizon 2020 project ENVRIPLUSintends to address this interoperability challenge as it relates to the design, implementation and operation of environmental science RIs, the theme focuses on key issues of data identification and citation, curation, cataloguing, processing, optimization, and provenance, supported by a generic cross-infrastructure reference model.
Zhiming Zhao, Paul Martin 0002, Paola Grosso, Wouter Los, Cees T. A. M. de Laat, Keith Jeffrey, Alex R. Hardisty, Alex Vermeulen, Donatella Castelli, Yannick Legré, Werner Kutsch
e-Science1
2013 An Autonomous Security Storage Solution for Data-Intensive Cooperative Cloud Computing
abstract
In order to reduce untrustworthy between cloud users and the underlying cloud storage platform, a novel cloud security storage solution is proposed based on autonomous data storage, management, and access control. The roles of users are re-evaluated, and the knowledge provided by the users is incorporated into the cloud storage model. Both the superiority of the public cloud in large scale data storage and the advantages of the private cloud in privacy preserving can be obtained. The main advantages of our approach include avoiding the superposition of complex security policies and overcoming the mistrust between the users and the platform. Furthermore, our security storage service can be easily integrated into the cooperative cloud computing environment. A prototype system is developed, and a use case is also presented.
Wenchao Jiang, Zhiming Zhao, Cees T. A. M. de Laat
e-Science2
2013 Dynamic Workflow Planning on Programmable Infrastructure
abstract
The Network Service Interface (NSI) has been created as a result of collaborative development of network and application engineers primarily associated with the Research and Education (R&E) community. The NSI allows workflow systems not only to check available service points for a workflow engine to schedule executions, but also to reserve and provide network connections among those service points. The Open Flow technology provides programmability on the network Flow and allows software to define dynamically behaviour of the network. These new features offer data intensive applications new opportunities to optimize the mapping between data Flow patterns and the infrastructure yielding better system level quality. However, they also require the computing support systems effectively capture not only the characteristics of the application workflow but also the controllability of the underlying network. In this paper we discussed the extension of our previous system called Network QoS Planner (NEWQoSPlanner) and investigated how reservation based connection services can be enhanced by dynamic network Flow control. We also discusse how NEWQoSPlanner invokes network services to achieve connection reservation and provisioning, and includes Open Flow to realize dynamic Flow optimization for data intensive workflows.
Wenchao Jiang, Zhiming Zhao, Adianto Wibisono, Paola Grosso, Cees T. A. M. de Laat
NAS2
2012 Addressing Big Data challenges for Scientific Data Infrastructure
abstract
This paper discusses the challenges that are imposed by Big Data Science on the modern and future Scientific Data Infrastructure (SDI). The paper refers to different scientific communities to define requirements on data management, access control and security. The paper introduces the Scientific Data Lifecycle Management (SDLM) model that includes all the major stages and reflects specifics in data management in modern e-Science. The paper proposes the SDI generic architecture model that provides a basis for building interoperable data or project centric SDI using modern technologies and best practices. The paper explains how the proposed models SDLM and SDI can be naturally implemented using modern cloud based infrastructure services provisioning model.
Yuri Demchenko, Zhiming Zhao, Paola Grosso, Adianto Wibisono, Cees T. A. M. de Laat
CloudCom2
2012 OEIRM: An Open Distributed Processing Based Interoperability Reference Model for e-Science
Zhiming Zhao, Paola Grosso, Cees T. A. M. de Laat
NPC1
2011 Resource Discovery in Large Scale Network Infrastructure
abstract
Semantic web technologies provide a standardised mechanism for describing and accessing the services of underlying infrastructure. These technologies facilitate the inclusion of the quality of network services in the control loop of high level applications and allow applications to tune the system level performance with additional quality dimensions. However, the descriptions of a large infrastructure are often composed and maintained by different parties and can have different levels of details because of the administration policies. These facts make the development of high level applications unnecessarily difficult. We present a preprocessing framework to hide these difficulties from high level application developers by transforming, integrating, and filtering raw descriptions of the infrastructure into proper information content that these applications need.
Zhiming Zhao, Arie Taal, Paola Grosso, Cees T. A. M. de Laat
NAS1
2009 Special section on workflow systems and applications in e-Science
Zhiming Zhao, Adam Belloum, Marian Bubak
Future Gener. Comput. Syst.1
2008 A Framework for Interactive Parameter Sweep Applications
abstract
Summary form only given. A typical parameter sweep application (PSA) is a parameterized application which has to be executed independently large number of times, to locate a particular point in the parameter space that satisfies certain criteria. From the perspective of domain scientists, the complexity of underlying grid environment should be hidden so that domain scientists can focus on their main concern on performing their experiments. The system needs to provide friendly user environment for scientists to change parameter space or set new policy for execution at runtime. The framework should provide flexible interface for porting legacy applications to a PSA. We proposed five functional components for the interactive framework: a GUI for user to describe experiment, a visualizer to presenting computing results, a coordinator to schedule the execution of computing tasks, a repository to collect computing results, and interface to job scheduling tools, e.g., from Grid. In the current prototype, a tuple space like workspace is used to maintain state of experiment. This prototype allows basic interactivity such as modification of experiment state during execution which demonstrated the capability to manipulate parameter sweep tasks at run time.
Adianto Wibisono, Zhiming Zhao, Adam Belloum, Marian Bubak
CCGRID2
2008 Support for Cooperative Experiments in VL-e: From Scientific Workflows to Knowledge Sharing
abstract
Recent advances in Internet and Grid technologies have greatly enhanced processes in scientific experiments; not only computing and data intensive tasks become feasible, but also large scale collaborations between resources and users are now possible. Scientific workflows encode intelligence of successful experiments and become important resources. Sharing these resources and allowing scientists to cooperate in one experiment are essential to promote the knowledge transfer among scientists and to accelerate the scientific achievements. In this demo, we present a solution developed in the Dutch Virtual Laboratory for e-Science project for supporting cooperative experiments.
Zhiming Zhao, Adam Belloum, Marian Bubak, Louis O. Hertzberger
eScience1
2007 Using Jade agent framework to prototype an e-Science workflow bus
abstract
Most of the existing scientific workflow management systems (SWMS) are driven by applications from specific domains and are developed in academic projects. It is challenging to introduce an existing SWMS to a new domain; not only the workflow model and description language do not easily fit in new problem domains, but also the unstable development state of existing systems does not provide all functionality required by the new applications and thus gives high risk for the development. Aggregating different workflow systems as one generic environment enables the sharing on both components and processes between experiments, and promotes the knowledge transfer between domains. A workflow bus approach is to integrate different e-science workflow engines via a software bus. In this paper, we present the basic idea of workflow bus, and discuss how Jade agent framework can be used to prototype the runtime infrastructure of a workflow bus.
Zhiming Zhao, Adam Belloum, Cees T. A. M. de Laat, Pieter W. Adriaans, Louis O. Hertzberger
CCGRID1
2006 VLE-WFBus: A Scientific Workflow Bus for Multi e-Science Domains
abstract
In e-Science, a Grid environment enables data and computing intensive tasks and provides a new supporting infrastructure for scientific experiments. Scientific workflow management systems (SWMS) hide the integration details among Grid resources and allow scientists to prototype an experimental computing system at a high level of abstraction. However, the development of an effective SWMS requires profound knowledge on both application domains and the network programming, and is often time consuming and domain specific. Integrating mature implementations of domain specific SWMS improves reusability of workflow resources and promotes a generic framework for different e-Science domains. In this paper, we discuss different options to derive a generic workflow management system from domain specific implementations, and propose a workflow bus based solution, called VLE-WFBus. Legacy SWMSs are wrapped as federated components and are loosely coupled as one workflow system via a runtime infrastructure. An agent based prototype is presented; the integration among different workflow management systems has been demonstrated.
Zhiming Zhao, Suresh Booms, Adam Belloum, Cees T. A. M. de Laat, Louis O. Hertzberger
e-Science1
2005 Rapid Prototyping of Complex Interactive Simulation Systems
abstract
By allowing human users to manipulate simulation models and steer their execution at run time, interactive simulation systems (ISS) are essential to realise problem solving environments (PSE) for studying complex problems that are difficult to investigate using conventional methodologies. However, the complexity of system development critically hampers the construction of ISSs; when a scientist explores a complex problem, he has to spend much of his effort on various implementation issues, instead of on the investigation of the experiment itself A layered framework for developing ISSs is crucial to hide the underlying development issues from scientists and to allow them to focus on the high-level behaviour of the system. In this paper we investigate a solution to the complexity issues in ISSs based on the separation of application logic control and system functionality. We implement a proof of concept architecture, called interactive simulation system conductor, based on high level architecture (HLA), and demonstrate that this architecture allows a scientist to quickly adapt the system to his needs in a rapid prototyping approach. A medical application for planning surgical operations is used as a test case.
Zhiming Zhao, G. Dick van Albada, Peter M. A. Sloot
ICECCS1
2005 Agent Technology and Scientific Workflow Management in an E-Science Environment
abstract
In e-science environments, scientific workflow management systems (SWMS) hide the integration details among grid resources and allow scientists to prototype an experimental computing system at a high level of abstraction. However, the development of an effective SWMS requires profound knowledge on both application domains and the network programming, and is often time consuming. Agent technologies provide suitable solutions to decompose the control intelligence of flow execution and to encapsulate distributed e-science resources. The work presented in this paper is conducted in the context of the Dutch Virtual Laboratory for e-science (VL-e) project. Agent technologies are proposed to realize generic workflow support.
Zhiming Zhao, Adam Belloum, Peter M. A. Sloot, Louis O. Hertzberger
ICTAI1