VLDB 2026 Research / reviewers in the wild / expert
Shiyong Lu
dblp:l/ShiyongLu
· DBLP profile ↗
63ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-7864-1815ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 22 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 17 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 2 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Integration of Blockchain Technology in Collaborative Scientific WorkflowsabstractBlockchain technology has emerged as a transformative force across various sectors, especially in enhancing collaborative scientific workflows. This paper delves into the unique challenges and opportunities associated with integrating blockchain into these workflows. We conduct a comprehensive analysis of recent literature to identify key themes, categorize different approaches, and assess the potential of blockchain to improve data integrity, provenance, and collaborative research efforts among diverse stakeholders. Through this exploration, we aim to provide an insightful overview of the current landscape and propose directions for future research focused on the role of blockchain in facilitating effective collaboration in scientific endeavors. Shiyong Lu, Junwen Liu, Yong Zhao 0009, Changxin Bai |
IEEE Big Data | 1 |
| 2023 | A Generic Efficient Scientific Workflow Engine for the Optimizations of Run-Time ExecutionabstractWorkflow has proven to be a highly effective computing model for a variety of scientific applications, offering flexible data types and unstructured parallelism that surpasses simple parallel execution models such as MapReduce. However, current workflow management systems in cloud computing environments experience unnecessary delays in task execution due to the separation of task execution and data transfer processes, which causes a child task to wait until all its predecessor tasks complete, rather than waiting only for necessary input data becoming ready. The goal of this paper is to eliminate the unnecessary delay of child tasks in a workflow, which is achieved through a new workflow engine architecture that separates workflow planner from workflow executor in the general framework of the DATAVIEW scientific workflow management system. This new engine architecture can be generalized and applied to other workflow systems. Our design integrates a new task release mechanism based on a data dependency model with the workflow executor of DATAVIEW. This approach enables prompt task launching once input data becomes available, instead of waiting for all predecessor tasks to finish. The architecture employs distributed algorithms for implementing the workflow executor and the task executors, performing various optimization on data movement, task movement, and communication among different subsystems. The experiments show that our new architecture based on the new task release model can significantly reduce overall execution time of a workflow in DATAVIEW. Changxin Bai, Junwen Liu, Anik Tahabilder, M. M. Imran, Shiyong Lu, Dunren Che |
SSE | 5 |
| 2023 | Infrastructure-level Support for GPU-Enabled Deep Learning in DATAVIEW
Junwen Liu, Ziyun Xiao, Shiyong Lu, Dunren Che, Ming Dong 0001, Changxin Bai |
Future Gener. Comput. Syst. | 3 |
| 2022 | Securing Big Data Scientific Workflows via Trusted Heterogeneous EnvironmentsabstractBig data workflow management systems (BDWMS)s have recently emerged as popular data analytics platforms to conduct large-scale data analytics in the cloud. However, the protection of data confidentiality and secure execution of workflow applications remains an important and challenging problem. Although a few data analytics systems, such as VC3 and Opaque, were developed to address security problems, they are limited to specific domains such as Map-Reduce-style and SQL query workflows. A generic secure framework for BDWMSs is still missing. In this article, we propose SecDATAVIEW, a distributed BDWMS that employs heterogeneous workers, such as Intel SGX and AMD SEV, to protect both workflow and workflow data execution, addressing three major security challenges: (1) Reducing the TCB size of the big data workflow management system in the untrusted cloud by leveraging the hardware-assisted TEE and software attestation; (2) Supporting Java-written workflow tasks to overcome the limitation of SGX’s lack of support for Java programs; and (3) Reducing the adverse impact of SGX enclave memory paging overhead through a “Hybrid” workflow task scheduling system that selectively deploys sensitive tasks to a mix of SGX and SEV worker nodes. Our experimental results show that SecDATAVIEW imposes moderate overhead on the workflow execution time. Saeid Mofrad, Ishtiaq Ahmed, Fengwei Zhang, Shiyong Lu, Ping Yang 0002, Heming Cui |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | A Security Framework for Scientific Workflow Provenance Access Control PoliciesabstractThe notion of collaborative scientific workflow is coined to address the increasing need for collaborative data analytics. In collaborative environments, access control policies are necessary for controlling the sharing of workflows, data products, and provenance information among collaborating parties. In particular, the protection of workflow provenance is critical because it often encodes the detailed protocol of a scientific experiment and carries the intellectual property of the respective stakeholders. In addition, since scientific workflows often evolve quickly, the corresponding access control policies for workflow provenance have to evolve as well. It is important to ensure that the evolution of workflow provenance access control policies maintain certain properties, in order to guarantee the correctness and performance of the corresponding policy enforcement. In this paper, we 1) propose a role-based access control model for scientific workflow provenance; 2) define three quality requirements for scientific workflow provenance access control policies - consistency, completeness, and conciseness; 3) develop a mechanism mapping from specifications of workflows to their counterparts in a provenance that preserves such quality properties, and 4) conduct a case study on a scientific workflow for autism behavioral data analysis that demonstrates the feasibility of our proposed analysis algorithms. Fahima Amin Bhuyan, Shiyong Lu, Robert G. Reynolds, Jia Zhang 0001, Ishtiaq Ahmed |
IEEE Trans. Serv. Comput. | 2 |
| 2021 | Deep-Learning-as-a-Workflow (DLaaW): An Innovative Approach to Enabling Deep Learning in Scientific WorkflowsabstractScientific workflow has become a popular cyberinfrastructure paradigm to accelerate scientific discoveries by enabling scientists to formalize and structure complex scientific processes. With the recent success of deep learning models in many scientific applications, there is a rising need for infrastructure-level support for deep learning technologies in scientific workflow cyberinfrastructures. However, current scientific workflow cyberinfrastructures and GPU-enabled deep learning frameworks are developed separately, neither alone can be a satisfactory choice. In this paper, We propose the Deep-Learning-as-a-Workflow approach in DATAVIEW, which for the first time incorporates native infrastructure level support for GPU-enabled deep learning in a scientific workflow management system and enables the fast training and execution of neural networks as workflows (NNWorkflows) leveraging various types of GPU resource configurations. Our experiments demonstrate the salient usability feature of DATAVIEW in providing seamless infrastructure-level support to both scientific and deep learning workflows in one system, while delivering competitive (better in most cases) learning efficiency compared to the conventional implementations based on Keras. Junwen Liu, Ziyun Xiao, Shiyong Lu, Dunren Che |
IEEE BigData | 3 |
| 2019 | Multi-Objective Resource Mapping and Allocation for Volunteer Cloud ComputingabstractAlthough Virtual Machine placement has been extensively researched in traditional Cloud Computing environments, it remains an open challenging problem for Volunteer Cloud Computing, which exhibits several divergent characteristics, including intermittent availability of nodes and unreliable infrastructure. In this paper, we model the Virtual Machine placement problem in Volunteer Cloud Computing as a bounded 0-1 multi-dimensional knapsack problem and develop three heuristic based algorithms to meet the objectives and constraints specific to Volunteer Cloud Computing. Empirical evidences on a real Volunteer Cloud Computing test-bed show the competitive performance results of these algorithms. Tessema Mengistu, Dunren Che, Shiyong Lu |
CLOUD | 3 |
| 2019 | SecDATAVIEW: a secure big data workflow management system for heterogeneous computing environmentsabstractBig data workflow management systems (BDWFMSs) have recently emerged as popular platforms to perform large-scale data analytics in the cloud. However, the protection of data confidentiality and secure execution of workflow applications remains an important and challenging problem. Although a few data analytics systems were developed to address this problem, they are limited to specific structures such as Map-Reduce-style workflows and SQL queries. This paper proposes SecDATAVIEW, a BDWFMS that leverages Intel Software Guard eXtensions (SGX) and AMD Secure Encrypted Virtualization (SEV) to develop a heterogeneous trusted execution environment for workflows. SecDATAVIEW aims to (1) provide the confidentiality and integrity of code and data for workflows running on public untrusted clouds, (2) minimize the TCB size for a BDWFMS, (3) enable the trade-off between security and performance for workflows, and (4) support the execution of Java-based workflow tasks in SGX. Our experimental results show that SecDATAVIEW imposes 1.69x to 2.62x overhead on workflow execution time on SGX worker nodes, 1.04x to 1.29x overhead on SEV worker nodes, and 1.20x to 1.43x overhead on a heterogeneous setting in which both SGX and SEV worker nodes are used. Saeid Mofrad, Ishtiaq Ahmed, Shiyong Lu, Ping Yang 0002, Heming Cui, Fengwei Zhang |
ACSAC | 3 |
| 2019 | RAMP: Real-Time Anomaly Detection in Scientific WorkflowsabstractResearch integrity is crucial to ensuring the trustworthiness of scientific discoveries. This work is aimed at detecting misbehaviors targeting scientific workflows, which are computing paradigms widely used to facilitate scientific collaborations across multiple geographically distributed research sites. We develop a new system called RAMP(Real-Time Aggregated Matrix Profile) for real-time anomaly detection in scientific workflow systems. RAMP builds upon an existing time series data analysis technique called Matrix Profile to detect anomalous distances among subsequences of event streams collected from scientific workflows in an online manner. Using an adaptive uncertainty function, the anomaly detection model is dynamically adjusted to prevent high false alarm rates. RAMP can incorporate user feedback on reported anomalies and modify model parameters to improve anomaly detection accuracy. Our experimental results from applying RAMP to the logs generated by DATAVIEW, a scientific workflow platform, show that RAMP is able to identify a varied range of anomalies with high accuracy for both interleaved and non-interleaved workflow executions in real time. Jerome Dinal Herath, Changxin Bai, Guanhua Yan, Ping Yang 0002, Shiyong Lu |
IEEE BigData | 5 |
| 2018 | Semi-Markov Process Based Reliability and Availability Prediction for Volunteer Cloud SystemsabstractAlternative cloud systems that depend on spare resources of volunteer computers are emerging, complementing the conventional data center based clouds. These alternative cloud systems have one characteristic in common - they do not rely on dedicated data centers to provide the cloud services. Delivering reliable cloud services over volatile donated computing resources of volunteer hosts is one of the main challenges in such opportunistic cloud systems that run over scavenged resources. This paper discusses the design of a multi-state semi-Markov process based availability and reliability prediction model for nodes in volunteer cloud computing systems. It also presents the implementation and integration of the model in a real volunteer cloud system, cuCloud. Moreover, the evaluation and empirical results obtained from the experimentation performed on real cuCloud system is discussed. Tessema Mengistu, Dunren Che, Abdulrahman Alahmadi, Shiyong Lu |
IEEE CLOUD | 4 |
| 2017 | Predicting efficacy of therapeutic services for autism spectrum disorder using scientific workflowsabstractEarly intervention in autism, although deemed as essential, has high variance in the outcome attained, partially due to complex interaction between multitude of factors and variables involved, and the lack of systematic study to untangle their influences in the outcome. Therefore, pairing set of interventions with an individual children to cater for their need remains highly challenging. From the perspective of parents, unknown factors emanate from their unfamiliarity with what interventions are out there and why. From the perspective of caregivers, it is critical to understand unique attributes of the individual children develop over time. There is a scarcity of exploration of interactions between attributes specific to a child, family characteristics and therapeutic, medical and educational services. In this research, we aim to bridge the gap. In this study, we identify predictive features pertaining to each individual child and how they interact responding to different interventions and services. We have studied temporal data and model improvement/regression outcomes at different timestamped milestones and overlayed a model to aid parents and caregivers in coming up with pragmatic intervention plan. We propose a scientific workflow to automate the modeling process and rely on DATAVIEW to guarantee computational reproducibility and data fidelity. We use data collected by SFARI dataset for evaluation. To the best of our knowledge, this is first-time amalgamation between the Autism Health informatics community and the Workflow community; and this is the first-time study that combines prediction methods applied on Autism Spectrum Disorder (ASD) Phenotype data to provide guidance to parents and caregivers. Fahima Amin Bhuyan, Shiyong Lu, Ishtiaq Ahmed, Jia Zhang 0001 |
IEEE BigData | 2 |
| 2016 | Scheduling big data workflows in the cloud under budget constraintsabstractBig data is fast becoming a ubiquitous term in both academia and industry and there is a strong need for new data-centric workflow tools and techniques to process and analyze large-scale complex datasets that are growing exponentially. On the other hand, the unbound resource leasing capability foreseen in the cloud facilitates data scientists to wring actionable insights from the data in a time and cost efficient manner. In the data-centric workflow environment, scheduling data processing tasks onto appropriate resources are often driven by the constraints provided by the users. Enforcing a constraint while executing the workflow in the cloud adds a new optimization challenge on how to meet the objective while satisfying the given constraint. In this paper, we propose a new Big dAta woRkflow schEduler uNder budgeT constraint known as BARENTS that supports high-performance workflow scheduling in a heterogeneous cloud computing environment with a single objective to minimize the workflow makespan under a provided budget constraint. Our case study and experiments show the competitive advantages of our proposed scheduler. The proposed BARENTS scheduler is implemented in a new release of DATA VIEW, one of the most usable big data workflow systems in the community. Aravind Mohan, Mahdi Ebrahimi, Shiyong Lu, Alexander Kotov 0001 |
IEEE BigData | 3 |
| 2016 | Feedback or Research: Separating Pre-purchase from Post-purchase Consumer Reviews
Alexander Kotov 0001, Aravind Mohan, Shiyong Lu, Paul M. Stieg |
ECIR | 4 |
| 2015 | TPS: A task placement strategy for big data workflowsabstractWorkflow makespan is the total execution time for running a workflow in the Cloud. The workflow makespan significantly depends on how the workflow tasks and datasets are allocated and placed in a distributed computing environment such as Clouds. Incorporating data and task allocation strategies to minimize makespan delivers significant benefits to scientific users in receiving their results in time. The main goal of a task placement algorithm is to minimize the total amount of data movement between virtual machines during the execution of the workflows. In this paper, we do the following: 1) formalize the task placement problem in big data workflows; 2) propose a task placement strategy (TPS) that considers both initial input datasets and intermediate datasets to calculate the dependency between workflow tasks; and 3) perform extensive experiments in the distributed environment to demonstrate that the proposed strategy provides an effective task distribution and placement tool. Mahdi Ebrahimi, Aravind Mohan, Shiyong Lu, Robert G. Reynolds |
IEEE BigData | 3 |
| 2015 | Parametric and Non-parametric User-aware Sentiment Topic ModelsabstractThe popularity of Web 2.0 has resulted in a large number of publicly available online consumer reviews created by a demographically diverse user base. Information about the authors of these reviews, such as age, gender and location, provided by many on-line consumer review platforms may allow companies to better understand the preferences of different market segments and improve their product design, manufacturing processes and marketing campaigns accordingly. However, previous work in sentiment analysis has largely ignored these additional user meta-data. To address this deficiency, in this paper, we propose parametric and non-parametric User-aware Sentiment Topic Models (USTM) that incorporate demographic information of review authors into topic modeling process in order to discover associations between market segments, topical aspects and sentiments. Qualitative examination of the topics discovered using USTM framework in the two datasets collected from popular online consumer review platforms as well as quantitative evaluation of the methods utilizing those topics for the tasks of review sentiment classification and user attribute prediction both indicate the utility of accounting for demographic information of review authors in opinion mining. Zaihan Yang, Alexander Kotov 0001, Aravind Mohan, Shiyong Lu |
SIGIR | 4 |
| 2015 | Enabling scalable scientific workflow management in the Cloud
Yong Zhao 0009, Youfu Li 0002, Ioan Raicu, Shiyong Lu, Wenhong Tian, Heng Liu 0004 |
Future Gener. Comput. Syst. | 4 |
| 2015 | Typetheoretic Approach to the Shimming Problem in Scientific WorkflowsabstractWhen composing Web services into scientific workflows, users often face the so-called shimming problem when connecting two related but incompatible components. The problem is addressed by inserting a special kind of adaptors, called shims, that perform appropriate data transformations to resolve data type inconsistencies. However, existing shimming techniques provide limited automation and burden users with having to define ontological mappings, generate data transformations, and even manually write shimming code. In addition, these approaches insert many visible shims that clutter workflow design and distract user's attention from functional components of the workflow. To address these issues, we 1) reduce the shimming problem to a runtime coercion problem in the theory of type systems, 2) propose a scientific workflow model and define the notion of well-typed workflows, 3) develop an algorithm to typecheck workflows, 4) design a function that inserts “invisible shims”, or runtime coercions into workflows, thereby solving the shimming problem for any well-typed workflow, 5) implement our automated shimming technique, including all the proposed algorithms, lambda calculus, type system, and translation functions in our VIEW system and present two case studies to validate our approach. Andrey Kashlev, Shiyong Lu, Artem Chebotko |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | A Service Framework for Scientific Workflow Management in the CloudabstractCloud computing is an emerging computing paradigm that can offer unprecedented scalability and resources on demand, and is getting more and more adoption in the science community, while scientific workflow management systems provide essential support such as management of data and task dependencies, job scheduling and execution, provenance tracking, etc., to scientific computing. As we are entering into a “big data” era, it is imperative to migrate scientific workflow management systems into the cloud to manage the ever increasing data scale and analysis complexity. We propose a reference service framework for integrating scientific workflow management systems into various cloud platforms, which consists of eight major components, including Cloud Workflow Management Service, Cloud Resource Manager, etc., and six interfaces between them. We also present a reference framework for the implementation of Cloud Resource Manager, which is responsible for the provisioning and management of virtual resources in the cloud. We discuss our implementation of the framework by integrating the Swift scientific workflow management system with the OpenNebula and Eucalyptus cloud platforms, and demonstrate the capability of the solution using a NASA MODIS image processing workflow and a production deployment on the Science@Guoshi network with support for the Montage image mosaic workflow. Yong Zhao 0009, Youfu Li 0002, Ioan Raicu, Shiyong Lu, Cui Lin, Wenhong Tian, Ruini Xue |
IEEE Trans. Serv. Comput. | 4 |
| 2014 | Adapting Medical Image Processing Tasks to a Scalable Scientific Workflow SystemabstractIn this paper, we present a web-based medical image processing scientific workflow system called DATAVIEW. This platform is implemented to satisfy the need of physicians to process their data by computer scientists and therefore, overcome the problem of delay for communicating between these two groups and speed up the processing which is very important in emergency conditions. For employing this workflow system, clinicians do not require to install any software tools and can create, save, share, reuse, and run their workflows only using a web browser without knowing the implementation details of the service. For employing the DATAVIEW workflow system for the medical image processing purpose, we extend this workflow system by integrating medical image software tools to this platform which is one of the challenges of this research. Also, in order to provide processing the data in parallel, we integrate high-end computing resource like grids to the system to speed up the processing, which is another challenge for this work. This is very important in medical imaging field because some software tools are very time consuming and in the urgent time, the physicians need to process multiple patients in parallel and get the results fast. As a case study, we choose epilepsy, which is one of the most common brain disorders. Hajar Hamidian, Shiyong Lu, Satyendra P. Rana, Farshad Fotouhi, Hamid Soltanian-Zadeh |
SERVICES | 2 |
| 2014 | Devising a Cloud Scientific Workflow Platform for Big DataabstractScientific workflow management systems (SWFMSs) are facing unprecedented challenges from big data deluge. As revising all the existing workflow applications to fit into Cloud computing paradigm is impractical, thus migrating SWFMSs into the Cloud to leverage the functionalities of both Cloud computing and SWFMSs may provide a viable approach to big data processing. In this paper, we first discuss the challenges for scientific workflow applications and the available solutions in details, and analyze the essential requirements for a scientific computing Cloud platform. Then we propose a service framework to normalize the integration of SWFMS with Cloud computing. Meanwhile, we also present our implementation experience based on the service Framework. At last, we set up a series of experiments to demonstrate the capability of our implementation and use a Montage Image Mosaic Workflow as a showcase of the implementation. Yong Zhao 0009, Youfu Li 0002, Shiyong Lu, Ioan Raicu, Cui Lin |
SERVICES | 3 |
| 2014 | Satisfiability Analysis of Workflows with Control-Flow Patterns and Authorization ConstraintsabstractWorkflow security has become increasingly important and challenging in today's open service world. While much research has been conducted on various security issues of workflow systems, the workflow satisfiability problem, which asks whether a set of users together can complete a workflow, is recently identified as an important research problem that needs more investigation. In this paper, we study the computational complexity of the problem along two directions: one is by considering either one path or all paths of a workflow, and the other is by considering the possible patterns in a workflow. We have shown that the general workflow satisfiability analysis problem is intractable. This result motivates us to consider restrictions on workflow control-flow patterns and access control policies, and to identify tractable cases of practical interest. Ping Yang 0002, Xing Xie 0002, Indrakshi Ray, Shiyong Lu |
IEEE Trans. Serv. Comput. | 4 |
| 2014 | Confucius: A Tool Supporting Collaborative Scientific Workflow CompositionabstractModern scientific data management and analysis usually rely on multiple scientists with diverse expertise. In recent years, such a collaborative effort is often structured and automated by a data flow-oriented process called scientific workflow. However, such workflows may have to be designed and revised among multiple scientists over a long time period. Existing workbenches are single user-oriented and do not support scientific workflow application development in a "collaborative fashion". In this paper, we report our research on the enabling techniques in the aspects of collaboration provenance management and reproduciability. Based on a scientific collaboration ontology, we propose a service-oriented collaboration model supported by a set of composable collaboration primitives and patterns. The collaboration protocols are then applied to support effective concurrency control in the process of collaborative workflow composition. We also report the design and development of Confucius, a service-oriented collaborative scientific workflow composition tool that extends an open-source, single-user development environment. Jia Zhang 0001, Daniel Kuc, Shiyong Lu |
IEEE Trans. Serv. Comput. | 3 |
| 2013 | Storing, Indexing and Querying Large Provenance Data Sets as RDF Graphs in Apache HBaseabstractProvenance, which records the history of an in-silico experiment, has been identified as an important requirement for scientific workflows to support scientific discovery reproducibility, result interpretation, and problem diagnosis. Large provenance datasets are composed of many smaller provenance graphs, each of which corresponds to a single workflow execution. In this work, we explore and address the challenge of efficient and scalable storage and querying of large collections of provenance graphs serialized as RDF graphs in an Apache HBase database. Specifically, we propose: (i) novel storage and indexing techniques for RDF data in HBase that are better suited for provenance datasets rather than generic RDF graphs and (ii) novel SPARQL query evaluation algorithms that solely rely on indices to compute expensive join operations, make use of numeric values that represent triple positions rather than actual triples, and eliminate the need for intermediate data transfers over a network. The empirical evaluation of our algorithms using provenance datasets and queries of the University of Texas Provenance Benchmark confirms that our approach is efficient and scalable. Artem Chebotko, John Abraham, Pearl Brazier, Anthony Piazza, Andrey Kashlev, Shiyong Lu |
SERVICES | 6 |
| 2013 | OPQL: Querying scientific workflow provenance at the graph level
Chunhyeok Lim, Shiyong Lu, Artem Chebotko, Farshad Fotouhi, Andrey Kashlev |
Data Knowl. Eng. | 2 |
| 2012 | A Dataflow-Based Scientific Workflow Composition FrameworkabstractScientific workflow has recently become an enabling technology to automate and speed up the scientific discovery process. Although several scientific workflow management systems (SWFMSs) have been developed, a formal scientific workflow composition model in which workflow constructs are fully compositional one with another is still missing. In this paper, we propose a dataflow-based scientific workflow composition framework consisting of (1) a dataflow-based scientific workflow model that separates the declaration of the workflow interface from the definition of its functional body; (2) a set of workflow constructs, including Map, Reduce, Tree, Loop, Conditional, and Curry, which are fully compositional one with another; (3) a dataflow-based exception handling approach to support hierarchical exception propagation and user-defined exception handling. Our workflow composition framework is unique in that workflows are the only operands for composition; in this way, our approach elegantly solves the two-world problem in existing composition frameworks, in which composition needs to deal with both the world of tasks and the world of workflows. The proposed framework is implemented and several case studies are conducted to validate our techniques. Xubo Fei, Shiyong Lu |
IEEE Trans. Serv. Comput. | 2 |
| 2011 | Scheduling Scientific Workflows Elastically for Cloud ComputingabstractMost existing workflow scheduling algorithms only consider a computing environment in which the number of compute resources is bounded. Compute resources in such an environment usually cannot be provisioned or released on demand of the size of a workflow, and these resources are not released to the environment until an execution of the workflow completes. To address the problem, we firstly formalize a model of a Cloud environment and a workflow graph representation for such an environment. Then, we propose the SHEFT workflow scheduling algorithm to schedule a workflow elastically on a Cloud computing environment. Our preliminary experiments show that SHEFT not only outperforms several representative workflow scheduling algorithms in optimizing workflow execution time, but also enables resources to scale elastically at runtime. Cui Lin, Shiyong Lu |
IEEE CLOUD | 2 |
| 2011 | Storing, reasoning, and querying OPM-compliant scientific workflow provenance using relational databases
Chunhyeok Lim, Shiyong Lu, Artem Chebotko, Farshad Fotouhi |
Future Gener. Comput. Syst. | 2 |
| 2011 | Data Replication in Data Intensive Scientific Applications with Performance GuaranteeabstractData replication has been well adopted in data intensive scientific applications to reduce data file transfer time and bandwidth consumption. However, the problem of data replication in Data Grids, an enabling technology for data intensive applications, has proven to be NP-hard and even non approximable, making this problem difficult to solve. Meanwhile, most of the previous research in this field is either theoretical investigation without practical consideration, or heuristics-based with little or no theoretical performance guarantee. In this paper, we propose a data replication algorithm that not only has a provable theoretical performance guarantee, but also can be implemented in a distributed and practical manner. Specifically, we design a polynomial time centralized replication algorithm that reduces the total data file access delay by at least half of that reduced by the optimal replication solution. Based on this centralized algorithm, we also design a distributed caching algorithm, which can be easily adopted in a distributed environment such as Data Grids. Extensive simulations are performed to validate the efficiency of our proposed algorithms. Using our own simulator, we show that our centralized replication algorithm performs comparably to the optimal algorithm and other intuitive heuristics under different network parameters. Using GridSim, a popular distributed Grid simulator, we demonstrate that the distributed caching technique significantly outperforms an existing popular file caching technique in Data Grids, and it is more scalable and adaptive to the dynamic change of file access patterns in Data Grids. Dharma Teja Nukarapu, Bin Tang 0004, Shiyong Lu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2010 | A Collectional Data Model for Scientific Workflow CompositionabstractModern scientific computations are usually data intensive, involving large-scale, heterogeneous and structured scientific datasets. Modeling, organizing, and processing scientific data have become key challenges for scientific workflow management systems (SWFMSs). In contrast to business data, which is usually relational and stored in databases, scientific data is often hierarchically organized and collection oriented. Although several data models have been proposed for SWFMSs, none of them provides a formal data model with a set of well-defined operators. In this paper, we take a first step towards formalizing a collection-oriented data model, called collectional data model, to model hierarchical collection oriented scientific data, and a set of well-defined operators to manipulate and query such data. We then apply the collectional data model to VIEW, a dataflow-based scientific workflow composition framework, whose workflow constructs are extended to support collections. We implement our techniques and validate them by a case study in a biological simulation project. Xubo Fei, Shiyong Lu |
ICWS | 2 |
| 2010 | Confucius: A Scientific Collaboration System Using Collaborative Scientific WorkflowsabstractLarge-scale scientific data management and analysis usually relies on many distributed scientists with diverse expertise. In recent years, such a collaborative effort is often composed and automated into a dataflow-oriented process, a so-called scientific workflow. However, existing scientific workflow tools are single user-oriented and do not support collaborative scientific workflow composition, execution, and management among multiple distributed scientists. In this paper, we report our study of collaboration protocols towards building a tool supporting collaborative scientific workflow composition. Based on a scientific collaboration ontology, we propose a collaboration model supported by a set of collaboration primitives and patterns. The collaboration protocols are then applied to support effective concurrency control in the process of collaborative workflow composition. Jia Zhang 0001, Daniel Kuc, Shiyong Lu |
ICWS | 3 |
| 2010 | RDFProv: A relational RDF store for querying and managing scientific workflow provenance
Artem Chebotko, Shiyong Lu, Xubo Fei, Farshad Fotouhi |
Data Knowl. Eng. | 2 |
| 2010 | Information flow analysis of scientific workflows
Ping Yang 0002, Shiyong Lu, Mikhail I. Gofman, Zijiang Yang 0006 |
J. Comput. Syst. Sci. | 2 |
| 2010 | Coclustering for cross-subject fiber tract analysis through diffusion tensor imagingabstractOne of the fundamental goals of computational neuroscience is the study of anatomical features that reflect the functional organization of the brain. The study of physical associations between neuronal structures and the examination of brain activity in vivo have given rise to the concept of anatomical and functional connectivity, which has been invaluable for our understanding of brain mechanisms and their plasticity during development. However, at present, there is no robust and accurate computational framework for the quantitative assessment of cortical connectivity patterns. In this paper, we present a quantitative analysis and modeling tool that is able to characterize anatomical connectivity patterns based on a newly developed coclustering algorithm, termed the business model-based coclustering algorithm (BCA). We apply BCA to diffusion tensor imaging (DTI) data in order to provide an automated and reproducible assessment of the connectivity patterns between different cortical areas in human brains. BCA not only partitions the cortical mantel into well-defined clusters, but also maximizes the connectivity strength between these clusters. Moreover, BCA is computationally robust and allows both outlier detection as well as parameter-independent determination of the number of clusters. Our coclustering results have showed good performance of BCA in identifying major white matter fiber bundles in human brains and facilitate the detection of abnormal connectivity patterns in patients suffering from various neurological diseases. Cui Lin, Darshan Pai, Shiyong Lu, Otto Muzik, Jing Hua 0001 |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2010 | Secure Abstraction Views for Scientific Workflow Provenance QueryingabstractProvenance has become increasingly important in scientific workflows and services computing to capture the derivation history of a data product, including the original data sources, intermediate data products, and the steps that were applied to produce the data product. In many cases, both scientific results and the used protocol are sensitive and effective access control mechanisms are essential to protect their confidentiality. In this paper, we propose: 1) a formal scientific workflow provenance model as the basis for querying and access control for workflow provenance; 2) a security model for fine-grained access control for multilevel provenance and an algorithm for the derivation of a full security specification based on inheritance, overriding, and conflict resolution; 3) a formalization of the notion of security views and an algorithm for security view derivation; and 4) a formalization of the notion of secure abstraction views and an algorithm for its computation. A prototype called SecProv has been developed, and experiments show the effectiveness and efficiency of our approach. Artem Chebotko, Shiyong Lu, Seunghan Chang, Farshad Fotouhi, Ping Yang 0002 |
IEEE Trans. Serv. Comput. | 2 |
| 2009 | A MapReduce-Enabled Scientific Workflow Composition FrameworkabstractMapReduce has recently gained a lot of attention as a parallel programming model for scalable data-intensive business and scientific analysis. In order to benefit from this powerful programming model in a scientific workflow environment, we propose a MapReduce-enabled scientific workflow composition framework consisting of: i) a dataflow based scientific workflow model that separates the declaration of the workflow interface from the definition of its functional body; ii) a set of dataflow constructs, including Map, Reduce, Loop, and Conditional, and their composition semantics to enable MapReduce-style scientific workflows; iii) an XML-based scientific workflow specification language, called WSL, in which both Map and Reduce are fully composable with other dataflow constructs in both flat and hierarchical manners. Besides leveraging the power of MapReduce to the workflow level, our workflow composition framework is unique in that workflows are the only operands for composition; in this way, our approach elegantly solves the two-world problem of existing composition frameworks, in which composition needs to deal with both the world of tasks and the world of workflows. The proposed framework is implemented and a case study is conducted to validate our techniques. Xubo Fei, Shiyong Lu, Cui Lin |
ICWS | 2 |
| 2009 | Collaborative Scientific WorkflowsabstractIn recent years, a number of scientific workflow management systems (SWFMSs) have been developed to help domain scientists synergistically integrate distributed computations, datasets, and analysis tools to enable and accelerate scientific discoveries. As more scientific research projects become collaborative in nature, there is a compelling need of dedicated services to support collaborative scientific workflows on the Internet. This paper reviews the state of the art of the field of scientific workflows towards the support of collaborative scientific workflows, identifies critical research challenges, and presents our ongoing research work aiming to study how to create services supporting collaborative scientific workflows. Shiyong Lu, Jia Zhang 0001 |
ICWS | 1 |
| 2009 | Semantics preserving SPARQL-to-SQL translation
Artem Chebotko, Shiyong Lu, Farshad Fotouhi |
Data Knowl. Eng. | 2 |
| 2009 | Atomicity and provenance support for pipelined scientific workflows
Shiyong Lu, Xubo Fei, Artem Chebotko, H. Victoria Bryant, Jeffrey L. Ram |
Future Gener. Comput. Syst. | 2 |
| 2009 | A Reference Architecture for Scientific Workflow Management Systems and the VIEW SOA SolutionabstractScientific workflows have recently emerged as a new paradigm for scientists to formalize and structure complex and distributed scientific processes to enable and accelerate many scientific discoveries. In contrast to business workflows, which are typically control flow oriented, scientific workflows tend to be dataflow oriented, introducing a new set of requirements for system development. These requirements demand a new architectural design for scientific workflow management systems (SWFMSs). Although several SWFMSs have been developed that provide much experience for future research and development, a study from an architectural perspective is still missing. The main contributions of this paper are: 1) based on a comprehensive survey of the literature and identification of key requirements for SWFMSs, we propose the first reference architecture for SWFMSs; 2) according to the reference architecture, we further propose a service-oriented architecture for View (a VIsual sciEntific Workflow management system); 3) we implemented View to validate the feasibility of the proposed architectures; and 4) we present a View-based scientific workflow application system (SWFAS), called FiberFlow, to showcase the application of our View system. Cui Lin, Shiyong Lu, Xubo Fei, Artem Chebotko, Darshan Pai, Zhaoqiang Lai, Farshad Fotouhi, Jing Hua 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2008 | Scientific Workflow Provenance Querying with Security ViewsabstractProvenance, the metadata that pertains to the derivation history of a data product, has become increasingly important in scientific workflow environments. In many cases, both data products and their provenance can be sensitive and effective access control mechanisms are essential to protect their confidentiality. In this paper, we propose i) a formalization of scientific workflow provenance as the basis for querying and access control; ii) a security specification mechanism for provenance at various granularity levels and the derivation of a full security specification based on inheritance, overriding, and conflict resolution rules; iii) a formalization of security views that are derived from a scientific workflow run provenance for different roles of users; and iv) a framework that integrates abstraction views and security views such that a user can examine provenance at different abstraction levels while respecting the security policy prescribed for her. We have developed the SecProv prototype to validate the effectiveness of our approach. Artem Chebotko, Seunghan Chang, Shiyong Lu, Farshad Fotouhi, Ping Yang 0002 |
WAIM | 3 |
| 2008 | Efficient Processing of RDF Queries with Nested Optional Graph Patterns in an RDBMSabstractRelational technology has shown to be very useful for scalable Semantic Web data management. Numerous researchers have proposed to use RDBMSs to store and query voluminous RDF data using SQL and RDF query languages. In this article, we study how RDF queries with the socalled well-designed graph patterns and nested optional patterns can be efficiently evaluated in an RDBMS. We propose to extend relational databases with a novel relational operator, nested optional join (NOJ), that is more efficient than left outer join in processing nested optional patterns of well-designed graph patterns. We design three efficient algorithms to implement the new operator in relational databases: (1) nested-loops NOJ algorithm (NL-NOJ); (2) sortmerge NOJ algorithm (SM-NOJ); and (3) simple hash NOJ algorithm (SH-NOJ). Based on a real-life RDF dataset, we demonstrate the efficiency of our algorithms by comparing them with the corresponding left outer join implementations and explore the effect of join selectivity on the performance of our algorithms. Artem Chebotko, Shiyong Lu, Mustafa Atay, Farshad Fotouhi |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2007 | XML-to-SQL Query Mapping in the Presence of Multi-valued Schema Mappings and Recursive XML Schemas
Mustafa Atay, Artem Chebotko, Shiyong Lu, Farshad Fotouhi |
DEXA | 3 |
| 2007 | Storing and Querying Scientific Workflow Provenance Metadata Using an RDBMSabstractProvenance management has become increasingly important to support scientific discovery reproducibility, result interpretation, and problem diagnosis in scientific workflow environments. This paper proposes an approach to provenance management that seamlessly integrates the interoperability, extensibility, and reasoning advantages of semantic Web technologies with the storage and querying power of an RDBMS. Specifically, we propose: i) two schema mapping algorithms to map an arbitrary OWL provenance ontology to a relational database schema that is optimized for common provenance queries; ii) two efficient data mapping algorithms to map provenance RDF metadata to relational data according to the generated relational database schema, and iii) a schema-independent SPARQL-to-SQL translation algorithm that is optimized on-the-fly by using the type information of an instance available from the input provenance ontology and the statistics of the sizes of the tables in the database. Experimental results are presented to show that our algorithms are efficient and scalable. Artem Chebotko, Xubo Fei, Cui Lin, Shiyong Lu, Farshad Fotouhi |
eScience | 4 |
| 2007 | Formal Modeling and Analysis of Scientific Workflows Using Hierarchical State MachinesabstractScientific workflows have recently emerged as a new paradigm for representing and managing complex distributed scientific computations and data analysis, and have enabled and accelerated many scientific discoveries. Many scientific workflows are distributed and collaborative as they result from some collaborative research projects that involve a number of geographically distributed organizations. In these workflows, information flow control becomes a key security problem. In this paper, we propose to model a scientific workflow using a hierarchical state machine and present techniques for verifying and controlling information propagation in scientific workflow environments based on hierarchical state machines. To the best of our knowledge, this is the first effort for information flow analysis in the area of scientific workflows. Ping Yang 0002, Zijiang Yang 0006, Shiyong Lu |
eScience | 3 |
| 2007 | GFBA: A Biclustering Algorithm for Discovering Value-Coherent Biclusters
Xubo Fei, Shiyong Lu, Horia F. Pop, Lily R. Liang |
ISBRA | 2 |
| 2007 | Coclustering Based Parcellation of Human Brain Cortex Using Diffusion Tensor MRI
Cui Lin, Shiyong Lu, Danqing Wu, Jing Hua 0001, Otto Muzik |
ISBRA | 2 |
| 2007 | XML subtree reconstruction from relational storage of XML documents
Artem Chebotko, Mustafa Atay, Shiyong Lu, Farshad Fotouhi |
Data Knowl. Eng. | 3 |
| 2007 | Efficient schema-based XML-to-Relational data mapping
Mustafa Atay, Artem Chebotko, Shiyong Lu, Farshad Fotouhi |
Inf. Syst. | 4 |
| 2006 | Runtime Security Verification for Itinerary-Driven Mobile AgentsabstractWe present a new approach to ensure the secure execution of itinerary-driven mobile agents, in which the specification of the navigational behavior of an agent is separated from the specification of its computational behavior. We empower each host with an access control policy so that the host will deny the access from an agent whose itinerary does not conform to the host's access control policy. A host uses model checking algorithms to check if the itinerary of the agent conforms to its access control policy written in mu-calculus, and if so, grant access permission. In order to address the state explosion problem for model checking itineraries, we propose an approach called model generation code. In this approach, instead of verifying the itinerary itself, a host actually checks the conservative models of a mobile agent. If a conservative model does not satisfy the host's access control policy, the mobile agent will provide refined models for further verification. Our preliminary results show that this is a practical and promising approach to ensure the secure execution of mobile agents Zijiang Yang 0006, Shiyong Lu, Ping Yang 0002 |
DASC | 2 |
| 2006 | Mining Correlation between Motifs and Gene ExpressionabstractOne of the major challenges in the post-genomic era is to determine all DNA-binding transcription factors (TFs) and their regulatory binding sites (motifs) within the genomes. To discover the relationship between the motifs and changes in gene expression, we propose a new algorithm, co-miner (correlation miner). Correlation rules are generated based on the expression profiles of genes with significant expression change through the time course of gene expression. Thus, we may consider the change in gene expression to be causatively associated with the transcription binding sites in the upstream sequences. In addition, we introduce partition and constraint pushing techniques to improve the performance and demonstrate their effectiveness by our experiments. By applying co-miner to a yeast dataset, the relationships between motifs and gene expression revealed by co-miner are confirmed in the literature. Yi Lu 0015, Shiyong Lu, Adrian E. Platts, Stephen A. Krawetz |
ICDM | 2 |
| 2006 | FM-test: a fuzzy-set-theory-based approach to differential gene expression data analysisabstractBACKGROUND: Microarray techniques have revolutionized genomic research by making it possible to monitor the expression of thousands of genes in parallel. As the amount of microarray data being produced is increasing at an exponential rate, there is a great demand for efficient and effective expression data analysis tools. Comparison of gene expression profiles of patients against those of normal counterpart people will enhance our understanding of a disease and identify leads for therapeutic intervention. RESULTS: In this paper, we propose an innovative approach, fuzzy membership test (FM-test), based on fuzzy set theory to identify disease associated genes from microarray gene expression profiles. A new concept of FM d-value is defined to quantify the divergence of two sets of values. We further analyze the asymptotic property of FM-test, and then establish the relationship between FM d-value and p-value. We applied FM-test to a diabetes expression dataset and a lung cancer expression dataset, respectively. Within the 10 significant genes identified in diabetes dataset, six of them have been confirmed to be associated with diabetes in the literature and one has been suggested by other researchers. Within the 10 significantly overexpressed genes identified in lung cancer data, most (eight) of them have been confirmed by the literatures which are related to the lung cancer. CONCLUSION: Our experiments on synthetic datasets show that FM-test is effective and robust. The results in diabetes and lung cancer datasets validated the effectiveness of FM-test. FM-test is implemented as a Web-based application and is available for free at http://database.cs.wayne.edu/bioinformatics. Lily R. Liang, Shiyong Lu, Xuena Wang, Yi Lu 0015, Vinay Mandal, Dorrelyn Patacsil |
BMC Bioinform. | 2 |
| 2006 | Automatic workflow verification and generation
Shiyong Lu, Arthur J. Bernstein, Philip M. Lewis |
Theor. Comput. Sci. | 1 |
| 2005 | On the consistency of XML DTDs
Shiyong Lu, Yezhou Sun, Mustafa Atay, Farshad Fotouhi |
Data Knowl. Eng. | 1 |
| 2005 | An Ontology-Based Multimedia Annotator for the Semantic Web of Language EngineeringabstractThe development of the Semantic Web, the next-generation Web, greatly relies on the availability of ontologies and powerful annotation tools. However, there is a lack of ontology-based annotation tools for linguistic multimedia data. Existing tools either lack ontology support or provide limited support for multimedia. To fill the gap, we present an ontology-based linguistic multimedia annotation tool, OntoELAN, which features: (1) the support for OWL ontologies; (2) the management of language profiles, which allow the user to choose a subset of ontological terms for annotation; (3) the management of ontological tiers, which can be annotated with language profile terms and, therefore, corresponding ontological terms; and (4) storing OntoELAN annotation documents in XML format based on multimedia and domain ontologies. To our best knowledge, OntoELAN is the first audio/video annotation tool in the linguistic domain that provides support for ontology-based annotation. It is expected that the availability of such a tool will greatly facilitate the creation of linguistic multimedia repositories as islands of the Semantic Web of language engineering. Artem Chebotko, Shiyong Lu, Farshad Fotouhi, Anthony Aristar |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2004 | A model for abstract process specification, verification and compositionabstractAn abstract business process contains a description the protocol that a business process engages in without revealing the internal computation of the process. This description provides the information necessary to compose the process with other Web services. BPEL supports this by providing distinct dialects for specifying abstract and executable processes. Unfortunately, BPEL does not prevent complex computations from being included in an abstract process. This complicates the protocol description, unnecessarily reveals implementation details, and makes it difficult to analyze correctness. We propose some restrictions on the data manipulation constructs that can be used in an abstract BPEL process. The restrictions permit a full description of a protocol while hiding computation. A restricted abstract process can easily be converted into an abstract BPEL process or expanded into an executable BPEL process. Based on these restrictions we propose a formal model for a business process and use it as the basis of an algorithm for demonstrating the process. We then sketch an algorithm for synthesizing a protocol based on a formal specification of its outcome and the tasks available for its construction. Ziyang Duan, Arthur J. Bernstein, Philip M. Lewis, Shiyong Lu |
ICSOC | 4 |
| 2004 | Semantics Based Verification and Synthesis of BPEL4WS Abstract ProcessesabstractWe introduce a logic model to formally specify the semantics of workflows and their composite tasks described as BPEL4WS abstract processes. Based on the model, we present a set of inference rules to deduce the strongest postcondition and weakest precondition of a workflow and demonstrate that automatic workflow verification is possible due to the restrictions on data manipulation in an abstract process. We then sketch an algorithm that automatically synthesizes a workflow given its specification and a task library. Ziyang Duan, Arthur J. Bernstein, Philip M. Lewis, Shiyong Lu |
ICWS | 4 |
| 2004 | Incremental genetic K-means algorithm and its application in gene expression data analysisabstractBACKGROUND: In recent years, clustering algorithms have been effectively applied in molecular biology for gene expression data analysis. With the help of clustering algorithms such as K-means, hierarchical clustering, SOM, etc, genes are partitioned into groups based on the similarity between their expression profiles. In this way, functionally related genes are identified. As the amount of laboratory data in molecular biology grows exponentially each year due to advanced technologies such as Microarray, new efficient and effective methods for clustering must be developed to process this growing amount of biological data. RESULTS: In this paper, we propose a new clustering algorithm, Incremental Genetic K-means Algorithm (IGKA). IGKA is an extension to our previously proposed clustering algorithm, the Fast Genetic K-means Algorithm (FGKA). IGKA outperforms FGKA when the mutation probability is small. The main idea of IGKA is to calculate the objective value Total Within-Cluster Variation (TWCV) and to cluster centroids incrementally whenever the mutation probability is small. IGKA inherits the salient feature of FGKA of always converging to the global optimum. C program is freely available at http://database.cs.wayne.edu/proj/FGKA/index.htm. CONCLUSIONS: Our experiments indicate that, while the IGKA algorithm has a convergence pattern similar to FGKA, it has a better time performance when the mutation probability decreases to some point. Finally, we used IGKA to cluster a yeast dataset and found that it increased the enrichment of genes of similar function within the cluster. Yi Lu 0015, Shiyong Lu, Farshad Fotouhi, Youping Deng, Susan J. Brown |
BMC Bioinform. | 2 |
| 2004 | Correct Execution of Transactions at Different Isolation LevelsabstractMany transaction processing applications execute at isolation levels lower than SERIALIZABLE in order to increase throughput and reduce response time. However, the resulting schedules might not be serializable and, hence, not necessarily correct. The semantics of a particular application determines whether that application will run correctly at a lower level and, in practice, it appears that many applications do. The decision to choose an isolation level at which to run an application and the analysis of the correctness of the resulting execution is usually done informally. We develop a formal technique to analyze and reason about the correctness of the execution of an application at isolation levels other than SERIALIZABLE. We use a new notion of correctness, semantic correctness, a criterion weaker than serializability, to investigate correctness. In particular, for each isolation level, we prove a condition under which the execution of transactions at that level will be semantically correct. In addition to the ANSI/ISO isolation levels of READ UNCOMMITTED, READ COMMITTED, and REPEATABLE READ, we also prove a condition for correct execution at the READ-COMMITTED with first-committer-wins and at SNAPSHOT isolation. We assume that different transactions in the same application can be executing at different levels, but that each transaction is executing at least at READ UNCOMMITTED. Shiyong Lu, Arthur J. Bernstein, Philip M. Lewis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Evolution based approaches to the preservation of endangered natural languagesabstractCultural algorithms, a form of evolutionary programming, employ a dual inheritance mechanism at population and knowledge levels to support problem solving, reasoning and knowledge extraction. Domain knowledge is extracted and separated from individuals within a population and is placed in a belief space. Hierarchical structures employed in the belief space help to accelerate and guide population evolution. The structure of the cultural algorithm lends itself well to a data rich, but knowledge poor distributed environment. In this paper we investigate the use of cultural algorithms to collect and mediate information collected from Web searches and Web services related to the task of acquiring and preserving knowledge about endangered languages where this knowledge about endangered languages where this knowledge is stored in a number of disparate sites. Jeffrey M. Stefan, Robert G. Reynolds, F. Fatouhi, Anthony Aristar, Shiyong Lu, Ming Dong 0001 |
IEEE Congress on Evolutionary Computation | 5 |
| 2003 | Conceptual Data Models for Engineering Information Modeling and Formal Transformation of EER and EXPRESS-G
Z. M. Ma, Shiyong Lu, Farshad Fotouhi |
ER | 2 |
| 2002 | A Structured Approach to Trade Negotiation Applications
Ziyang Duan, Albert Loo, Biswajit Sarkar, Shiyong Lu, Mark Van Loon, Subhra Bose |
CAINE | 4 |
| 2000 | Semantic Conditions for Correctness at Different Isolation LevelsabstractMany transaction processing applications execute at isolation levels lower than serializable in order to increase throughput and reduce response time. The problem is that non-serializable schedules are not guaranteed to be correct for all applications. The semantics of a particular application determines whether that application will run correctly at a lower isolation level, and in practice it appears that many applications do. Unfortunately, we know of an analysis technique that has been developed to test an application for its correctness at a particular level. Apparently decisions of this nature are made on an informal basis. In this paper we describe such a technique in a formal way. We use a new definition of correctness, semantic correctness, which is weaker than serializability, to investigate the correctness of such executions. For each isolation level, we prove a condition under which transactions that execute at that level will be semantically correct. In addition to the ANSI/ISO isolation levels of read uncommitted, read committed, and repeatable read, we also prove a condition for correct execution at the read committed with first-committer-wins (a variation of read committed) and at the snapshot isolation level. We assume that different transactions can be executing at different isolation levels, but that each transaction is executing at least at the read uncommitted level. Arthur J. Bernstein, Philip M. Lewis, Shiyong Lu |
ICDE | 3 |
| 1999 | Model Checking the Secure Electronic Transaction (SET) ProtocolabstractWe use model checking to establish five essential correctness properties of the secure electronic transaction (SET) protocol. SET has been developed jointly by Visa and MasterCard as a method to secure payment card transactions over open networks, and industrial interest in the protocol is high. Our main contributions are to firstly create a formal model of the protocol capturing the purchase request, payment authorization, and payment capture transactions. Together these transactions constitute the kernel of the protocol. We then encoded our model and the aforementioned correctness properties in the input language of the FDR model checker. Running FDR on this input established that our model of the SET protocol satisfies all five properties even though the cardholder and merchant, two of the participants in the protocol, may try to behave dishonestly in certain ways. To our knowledge, this is the first attempt to formalize the SET protocol for the purpose of model checking. Shiyong Lu, Scott A. Smolka |
MASCOTS | 1 |