VLDB 2026 Research / reviewers in the wild / expert
Beth Plale
dblp:54/6437 · also Beth A. Plale
· DBLP profile ↗
70ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0003-2164-8132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 10 first-author · 2 since 2021Software engineering, systems software and programming languages · 24 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Patra-RGCN: Missing Link Prediction in Model Card Graphs through Node Property EncodingsabstractDistributed frameworks for edgeAI span the edge to cloud continuum in support of AI through services for deployment, monitoring, and optimization. However, edge networks are unreliable, raising monitoring issues including missing events, network interruptions, and storage limitations all of which can lead to information loss that undermines decision-making and analytics. We address information loss through Patra Relational Graph Convolutional Networks (Patra-RGCN), an extension of Relational Graph Convolutional Networks (R-GCN), that uses both structural and content information about a graph (including time and Model Card subgraphs) to infer missing edges and suggest new connections. Experimental evaluation demonstrates that by effectively capturing both structural and attribute-level information, the proposed model significantly improves the detection of information loss in heterogeneous graphs. Krishna Priya, Sachith Withana, Beth Plale |
eScience | 3 |
| 2024 | Patra ModelCards: AI/ML Accountability in the Edge-Cloud ContinuumabstractThis paper introduces a framework for Model Cards, Patra ModelCards, that embeds model cards in the edge-cloud continuum for semi-automated information capture with the objective of greater trustworthiness and accountability for AI/ML models. Information captured includes fairness, explainability, and behavior of a model in different deployed environments. Our evaluation is of the framework’s claim of greater accountability. We evaluate the use of embedded vectors and similarity analysis to distinguish between deployed models that are duplicates of each other from those that represent a revision of an earlier developed model. The evaluation shows promising outcomes and good performance. Sachith Withana, Beth Plale |
e-Science | 2 |
| 2023 | CCGRID 2023: A Holistic Approach to Inclusion and Belongingabstract“CCGRID will act with responsibility as its primary consideration; with equity, diversity, and inclusion as its central goals.” from the CCGRID 2023 web site [1] Beth Plale, Preeti Malakar, Meenakshi D'Souza, Hemangee K. Kapoor, Yogesh L. Simmhan, Ilkay Altintas, S. Manohar 0001 |
CCGrid | 1 |
| 2023 | Democratization of AI: Challenges of AI Cyberinfrastructure and Software ResearchabstractLarge scale virtual teams working in the academic setting who assemble to carry out research that results in contributions to national cyberinfrastructure face a set of competing challenges. For the research team that is exploring AI, these challenges come into high relief when examined through the lense of responsibility to society. What form does responsibility for downstream uses of AI take in an NSF funded AI Institute, especially one that has as its priority to advance AI in and for cyberinfrastructure? In this brief abstract we identify the competing interests, and show how they raise challenges when analyzed through the lense of Democratizing AI. Our background is our now 2-year involvement in the NSF AI Institute Intelligent Cyberinfrastructure with Computational Learning in the Environment (ICICLE) [2]. Beth Plale, Sadia Khan, Alfonso Morales |
e-Science | 1 |
| 2023 | CKN: An Edge AI Distributed FrameworkabstractThe edge-cloud-HPC continuum is transformative for AI processing at the edge. With greater availability of both edge and cloud resources, AI inference, training, and optimization can be distributed across the continuum. We target edge-cloud in particular where the workload at the Edge server can exhibit discrete changes, for instance, when motion is detected. We optimize for Quality of Experience (QoE) and utilize historical data from the Edge, graphs, and Deep Learning to infer the next action to take. Using a large synthetic workload and publicly profiled inference models, our results show that predictive guidance outperforms random choice or best guess in optimal QoE of the edge-cloud continuum. Sachith Withana, Beth Plale |
e-Science | 2 |
| 2021 | Towards System for Knowledge Representation of Campaign ExperimentationabstractThe campaign is an experimentation construct for codesign activity wherein multiple researchers carry out computational experiments that individually contribute to a shared goal. The larger objective of our research is a system that exists in the experimental environment that constructs a knowledge representation of campaigns and products both produced and consumed such that the campaign can as efficient as possible and the products richly contextualized for reuse. Using campaign experiments running on the Summit machine at Oak Ridge National Labs, we demonstrate early results of support for discovery queries and for detecting when two sweeps are similar. Sachith Withana, Kshitij Mehta, Matthew Wolf, Beth Plale |
e-Science | 4 |
| 2021 | Transparency and Reproducibility Practice in Large-Scale Computational Science: A Preface to the Special SectionabstractWith this special section we bring you a practice and experience effort in transparency and reproducibility for large-scale computational science. A unique section, it consists of a research work plus six critques, each by a student team that reproduced the work. The original research work has been expanded in its science and also in its contribution to open science with a discussion of the student effort. Our letter contemplates implications as well. Beth Plale, Stephen Lien Harrell |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Reliable access to massive restricted texts: Experience-based evaluationabstractSummary Libraries are seeing growing numbers of digitized textual corpora that frequently come with restrictions on their content. Computational analysis corpora that are large, while of interest to scholars, can be cumbersome because of the combination of size, granularity of access, and access restrictions. Efficient management of such a collection for general access especially under failures depends on the primary storage system. In this paper, we identify the requirements of managing for computational analysis a massive text corpus and use it as basis to evaluate candidate storage solutions. The study based on the 5.9 billion page collection of the HathiTrust digital library. Our findings led to the choice of Cassandra 3.x for the primary back end store, which is currently in deployment in the HathiTrust Research Center. Zong Peng, Beth Plale |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Transparency by Design in eScience ResearchabstractBoth the landscape of eScience research and the environment in which the research is conducted are undergoing change. Transparency by design in eScience is proposed as a term to describe transparency in eScience practices, processes, methodologies, and research results. We break down different aspects of transparency and urge the eScience community towards a renewed commitment to scientific rigor because of the important role that we as scientists have to improve society and protect the good will that society has bestowed on science. Beth Plale |
eScience | 1 |
| 2018 | Big Provenance Stream Processing for Data Intensive ComputationsabstractIn the business and research landscape of today, data analysis consumes public and proprietary data from numerous sources, and utilizes any one or more of popular data-parallel frameworks such as Hadoop, Spark and Flink. In the Data Lake setting these frameworks co-exist. Our earlier work has shown that data provenance in Data Lakes can aid with both traceability and management. The sheer volume of fine-grained provenance generated in a multi-framework application motivates the need for on-the-fly provenance processing. We introduce a new parallel stream processing algorithm that reduces fine-grained provenance while preserving backward and forward provenance. The algorithm is resilient to provenance events arriving out-of-order. It is evaluated using several strategies for partitioning a provenance stream. The evaluation shows that the parallel algorithm performs well in processing out-of-order provenance streams, with good scalability and accuracy. Isuru Suriarachchi, Sachith Withana, Beth Plale |
eScience | 3 |
| 2017 | Pacific Rim Applications and Grid Middleware Assembly (PRAGMA): International clouds for data scienceabstractThis special issue presents a selection of research emerging from the Pacific Rim Applications and Grid Middleware Assembly (PRAGMA), an assembly of center-scale organizations around the Pacific Rim with membership from Australia, China, India, Indonesia, Japan, Malaysia, Philippines, South Korea, Thailand, Taiwan, the United States (California, Florida, Indiana, Virginia, and Wisconsin), and Vietnam.PRAGMA came into being in the early 2000s and through leveraging its unique makeup of widely distributed high-performance computing centers around the Pacific Rim, achieved advancements in grid computing then cloud computing and computer networks.In recent years, PRAGMA built on its foundation to advance scientific applications through focused scientific expeditions.In the process of advancing science through leading edge computational and networking capability, PRAGMA performs innovative research, education, and training, thus contributing to a better world through the human bonds of advancing science internationally.This special issue is an assemblage of the best papers appearing in the 2015 PRAGMA International Clouds for Data Science 2015 workshop held in conjunction with PRAGMA's 29th workshop, hosted by the University of Indonesia on its campus in Depok, Indonesia, in October 2015.The objective of the PRAGMA-ICDS 2015 workshop was to serve as a venue for presenting the latest research on the design, implementation, evaluation, and use of cloud technology, networking, and data management that enables new forms of research able to span international boundaries.The workshop received 10 papers, of which 7 were invited for further development for this special issue.All papers presented at the workshop appear in an open access print archive such as arXiv.orgor figshare.To mark the unique nature of this collection of papers and set the papers in a historical and collaborative context, the first paper in this collection is from Peter Arzberger. 1 Dr Arzberger founded the PRAGMA assembly in 2003 and has served as its chair and senior statesman since its inception and has been a significant influence on the assembly's culture of innovation and shared sense of trust.Dr Arzberger's article is a reflection on the 13 years of PRAGMA's accomplishment and a musing on opportunities and challenges for PRAGMA's future.When examining the nature of the 7 technical papers that make up this collection, one can see several themes emerging, the first of which is research that takes advantage of the unique international computer network of connected local clusters that connect the PRAGMA institutions.Many of the institutions of PRAGMA donate Beth Plale |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | KVLight: A Lightweight Key-Value Store for Distributed Access in CloudabstractKey-value stores (KVS) are finding use in Big Data applications as the store offers a flexible data model, scalability in number of distributed nodes, and high availability. In a cloud environment, a distributed KVS is often deployed over the local file system of the nodes in a cluster of virtual machines (VMs). Parallel file system (PFS) offers an alternate approach to disk storage, however a distributed key value store running over a parallel file system can experience overheads due to its unawareness of the PFS. Additionally, distributed KVS requires persistent running services which is not cost effective under the pay-as-you-go model of cloud computing because resources have to be held even under periods of no workload. We propose KVLight, a lightweight KVS that runs over PFS. It is lightweight in the sense that it shifts the responsibility of reliable data storage to the PFS and focuses on performance. Specifically, KVLight is built on an embedded KVS for high performance but uses novel data structures to support concurrent writes, giving capability that embedded KVSs are not currently designed for. Furthermore, it allows on-demand access without running persistent services in front of the file system. Empirical results show that KVLight outperforms Cassandra and Voldemort, two state-of-the-art KVSs, under both synthetic and realistic workloads. Jiaan Zeng, Beth Plale |
CCGrid | 2 |
| 2016 | Horme: Random Access Big Data AnalyticsabstractMapReduce is a parallel framework which has been widely adopted for conducting large-scale data analytics. In cases where analysis of multiple millions of books must be analyzed using federally funded high performance computing (HPC) resources, the framework fails to port directly. We propose a solution that builds off of MapReduce for use on a HPC system that preserves the key-value semantics of map-reduce while supporting the random access of query access for subsetting Big Data datasets, and at same time hosting the service using the storage medium found in HPC architectures (parallel file systems) for reduced latencies. Experimental results demonstrate Horme's good performance in the HPC setting, with up to 41.4% faster than NoSQL based solution in random access scenario. Guangchen Ruan, Beth Plale |
CLUSTER | 2 |
| 2016 | A hybrid approach to population construction for agricultural agent-based simulationabstractAn Agent Based Model (ABM) is a powerful tool for its ability to represent heterogeneous agents which through their interactions can reveal emergent phenomena. For this to occur though, the set of agents in an ABM has to accurately model a real world population to reflect its heterogeneity. But when studying human behavior in less well developed settings, the availability of the real population data can be limited, making it impossible to create agents directly from the real population. In this paper, we propose a hybrid method to deal with this data scarcity: we first use the available real population data as the baseline to preserve the true heterogeneity, and fill in the missing characteristics based on survey and remote sensing datasets; then for the remaining undetermined agent characteristics, we use the Microbial Genetic Algorithm to search for a set of values that can optimize the replicative validity of the model to match data observed from real world. We apply our method to the creation of a synthetic population of household agents for the simulation of agricultural decision making processes in rural Zambia. The result shows that the synthetic population created from the farmer register can correctly reflect the marginal distributions and the randomness of survey data; and can minimize the difference between the distribution of simulated yield and that of the observed yield in Post Harvest Survey (PHS). Peng Chen 0015, Tom Evans, Michael Frisby, Eduardo Izquierdo-Torres, Beth Plale |
eScience | 5 |
| 2016 | Crossing analytics systems: A case for integrated provenance in data lakesabstractThe volumes of data in Big Data, their variety and unstructured nature, have had researchers looking beyond the data warehouse. The data warehouse, among other features, requires mapping data to a schema upon ingest, an approach seen as inflexible for the massive variety of Big Data. The Data Lake is emerging as an alternate solution for storing data of widely divergent types and scales. Designed for high flexibility, the Data Lake follows a schema-on-read philosophy and data transformations are assumed to be performed within the Data Lake. During its lifecycle in a Data Lake, a data product may undergo numerous transformations performed by any number of Big Data processing engines leading to questions of traceability. In this paper we argue that provenance contributes to easier data management and traceability within a Data Lake infrastructure. We discuss the challenges in provenance integration in a Data Lake and propose a reference architecture to overcome the challenges. We evaluate our architecture through a prototype implementation built using our distributed provenance collection tools. Isuru Suriarachchi, Beth Plale |
eScience | 2 |
| 2016 | Argus: A Multi-tenancy NoSQL store with workload-aware resource reservation
Jiaan Zeng, Beth Plale |
Parallel Comput. | 2 |
| 2015 | ProvErr: System Level Statistical Fault Diagnosis Using Dependency ModelabstractLarge-scale distributed systems are difficult to debug in the event of failure. Yet rapid fault diagnosis that pinpoints failures to the component level is critical to fast recovery. We introduce a statistical approach to fault diagnosis that utilizes a dependency graph of execution to automatically discover the most probable fault cause(s) at a component level (either software or hardware resource). This approach leverages engineers' high level understanding of the system and requires a very small amount of information compared to existing methods. It also utilizes dependency information to eliminate redundant causes while retaining co-causes. Experiments using Apache Pig show that our approach has good, robust performance for diagnosing software bugs and resource shortages, and scales nearly linearly as system size increases. Peng Chen 0015, Beth Plale |
CCGRID | 2 |
| 2015 | Big Data Provenance Analysis and VisualizationabstractProvenance captured from E-Science experimentation is often large and complex, for instance, from agent-based simulations that have tens of thousands of heterogeneous components interacting over extended time periods. The subject of study of my dissertation is the use of E-Science provenance at scale. My initial research studied the visualization of large provenance graphs and proposed an abstract representation of provenance that supports useful data mining. Recent work involves analyzing large provenance data generated from agent-based simulations on a single machine. In continuation, I propose stream processing techniques to support the continuous and real-time analysis of data provenance, which is captured from agent based simulations on HPC and thus has unprecedented volume and complexity. Peng Chen 0015, Beth Plale |
CCGRID | 2 |
| 2015 | Workload-Aware Resource Reservation for Multi-tenant NoSQLabstractCloud hosted NoSQL data stores are for economic reasons often shared amongst multiple tenants simultaneously. The NoSQL provider consolidates multiple tenants access into a shared NoSQL instance and provides a dedicated view for each tenant. This multi-tenancy has tenants' data and workloads coexisting in the same node, which under certain conditions can lead to performance degradation of one tenant caused by another. In this paper, we investigate the multi-tenant interference in a common NoSQL store, HBase, and propose a resource reservation framework that reserves resources for prevention and dynamically adjusts the reservations according to tenant resource demands. The framework enforces cache reservation by splitting the cache space and disk reservation by scheduling requests to a distributed file system (DFS). A stochastic hill climbing algorithm is used to find a near-optimum plan for different resources reservations. Empirical results show that the framework can prevent interference and adapt to dynamic workloads under multi-tenancy. Jiaan Zeng, Beth Plale |
CLUSTER | 2 |
| 2015 | Towards Building a Lightweight Key-Value Store on Parallel File SystemabstractAs data grows in number and size, big data applications begin to revolutionize the underlying storage system. On one hand, key-value store has prevailed as the back-end storage for big data applications owning to its schema-less data model, high scalability, and etc. On the other hand, parallel file system shared by multiple nodes offers large-capacity, high-throughput, as well as high-bandwidth access and is used widely in high performance computing (HPC) and cloud computing environments. In this paper, we explore the opportunity of building a lightweight key-value store that supports concurrent access over a parallel file system. The key-value store proposed relies on the sharing nature of parallel file system to provide distributed access. Instead of organizing a cluster of nodes with long running services to delegate the access, our key-value store simply embeds itself into applications and requires no long running services neither communication between nodes. Such a design not only simplifies the structure of a distributed key-value store but also avoids overhead introduced by having running services around the file system. We implemented a prototype of this system and compared it against Cassandra, a state-of-art key-value store. Preliminary results are promising. Jiaan Zeng, Beth Plale |
CLUSTER | 2 |
| 2015 | Towards Sustainable Curation and Preservation: The SEAD Project's Data Services ApproachabstractWhen the effort to curate and preserve data is made at the end of a project, there is little opportunity to leverage ongoing research work to reduce curation costs or conversely, to leverage curation efforts to improve research productivity. In the Sustainable Environment Actionable Data (SEAD) project, we have envisioned a more active approach to data curation and preservation in which these processes occur in parallel with research and generate sufficient short and long-term return on researcher investments for self-interest to drive their adoption. In this paper, we describe the conceptual framework motivating the SEAD project and the suite of data services we have developed and deployed as an initial implementation of this approach. Use cases in which these services can reduce curation effort and aid ongoing research are highlighted and, based on our experience to date, we identify some key architectural features of our approach as well as open challenges to fully realizing the value of this approach in the broad ecosystem of cyberinfrastructure. James D. Myers, Margaret L. Hedstrom, Dharma Akmon, Sandra Payette, Beth Plale, Inna Kouper, D. Scott McCaulay, Robert H. McDonald, Isuru Suriarachchi, Aravindh Varadharaju, Praveen Kumar 0002, Mostafa Elag, Jong Lee, Rob Kooper, Luigi Marini |
e-Science | 5 |
| 2014 | Parallel and quantitative sequential pattern mining for large-scale interval-based temporal dataabstractMining frequent subsequences of patterns, or sequential pattern mining, has wide application in customer shopping sequence analysis, web log stream analysis, multi-modal behavioral studies, to name a few. To detect unknown, anomalous, and unexpected patterns from large-scale interval-based temporal data without complete a priori knowledge is challenging. In this paper, we present a framework - PESMiner which allows parallel and quantitative mining of sequential patterns at scale. Whereas most existing sequential mining algorithms can only find sequential orders of temporal events, our work presents a novel interactive temporal data mining algorithm capable of extracting precise temporal properties of sequential patterns. Furthermore, our work provides a unified parallel solution that scales our algorithms to larger temporal data sets by exploiting iterative MapReduce tasks. Comprehensive performance evaluations demonstrate that PESMiner significantly outperforms existing interval-based mining algorithms in terms of both quality (i.e. accuracy, precision, and recall) and scalability. Guangchen Ruan, Hui Zhang 0006, Beth Plale |
IEEE BigData | 3 |
| 2014 | Multi-tenant fair share in NoSQL data storesabstractNoSQL data stores see considerable attention today in big data, cloud hosted environments because of their fault tolerance, distribution and high availability. Shared NoSQL data stores are preferred for their ability to serve multiple tenants simultaneously which can improve resource utilization and lower management costs. Fair share in this setting can be a problem in that NoSQL data stores can be weak in preventing interference between tenants. We propose a methodology for multi-tenant fair share in a NoSQL store, in particular Cassandra. The approach uses an extended version of the deficit round robin algorithm to schedule tenant requests, and has local weight adjustment and slow tenant handling to improve the system throughput. Empirical results show that our approach is able to provide fair share for multi-tenancy. Jiaan Zeng, Beth Plale |
CLUSTER | 2 |
| 2014 | Study in Usefulness of Middleware-Only ProvenanceabstractData provenance is the lineage of a digital artifact or object. Its capture in workflow-controlled distributed applications is well studied but less is known about quality of provenance captured solely through existing control infrastructures (i.e., middleware frameworks used for high throughput computing). We study completeness of provenance in case where information is only available from the middleware layer. We use WorkQueue to validate our model. Our evaluation shows that provenance captured from a middleware framework is sufficient to represent the existence of output data and trace certain failures independent of the application semantics. We show the method's limitations as well. Devarshi Ghoshal, Beth Plale |
eScience | 3 |
| 2014 | Hierarchical MapReduce: towards simplified cross-domain data processingabstractSUMMARY The MapReduce programming model has proven useful for data‐driven high throughput applications. However, the conventional MapReduce model limits itself to scheduling jobs within a single cluster. As job sizes become larger, single‐cluster solutions grow increasingly inadequate. We present a hierarchical MapReduce framework that utilizes computation resources from multiple clusters simultaneously to run MapReduce job across them. The applications implemented in this framework adopt theMap–Reduce–GlobalReducemodel where computations are expressed as three functions: Map, Reduce, and GlobalReduce. Two scheduling algorithms are proposed, one that targets compute‐intensive jobs and another data‐intensive jobs, evaluated using a life science application, AutoDock, and a simple Grep. Data management is explored through analysis of the Gfarm file system.Copyright © 2012 John Wiley & Sons, Ltd. Beth Plale, Zhenhua Guo 0004, Wilfred W. Li, Judy Qiu, Yiming Sun 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Temporal representation for mining scientific data provenance
Peng Chen 0015, Beth Plale, Mehmet S. Aktas |
Future Gener. Comput. Syst. | 2 |
| 2013 | Dependency Provenance in Agent Based ModelingabstractResearchers who use agent-based models (ABM) to model social patterns often focus on the model's aggregate phenomena. However, aggregation of individuals complicates the understanding of agent interactions and the uniqueness of individuals. We develop a method for tracing and capturing the provenance of individuals and their interactions in the Net Logo ABM, and from this create a "dependency provenance slice", which combines a data slice and a program slice to yield insights into the cause-effect relations among system behaviors. To cope with the large volume of fine-grained provenance traces, we propose use-inspired filters to reduce the amount of provenance, and a provenance slicing technique called "non-preprocessing provenance slicing" that directly queries over provenance traces without recovering all provenance entities and dependencies beforehand. We evaluate performance and utility using a well known ecological Net Logo model called "wolf-sheep-predation". Peng Chen 0015, Beth Plale, Tom Evans |
e-Science | 2 |
| 2013 | Data Pipeline in MapReduceabstractMapReduce is an effective programming model for large scale text and data analysis. Traditional MapReduce implementation, e.g., Hadoop, has the restriction that before any analysis can take place, the entire input dataset must be loaded into the cluster. This can introduce sizable latency when the data set is large, and when it is not possible to load the data once, and process many times - a situation that exists for log files, health records and protected texts for instance. We propose a data pipeline approach to hide data upload latency in MapReduce analysis. Our implementation, which is based on Hadoop MapReduce, is completely transparent to user. It introduces a distributed concurrency queue to coordinate data block allocation and synchronization so as to overlap data upload and execution. The paper overcomes two challenges: a fixed number of maps scheduling and dynamic number of maps scheduling allows for better handling of input data sets of unknown size. We also employ delay scheduler to achieve data locality for data pipeline. The evaluation of the solution on different applications on real world data sets shows that our approach shows performance gains. Jiaan Zeng, Beth Plale |
e-Science | 2 |
| 2013 | Provenance Capture and Use in a Satellite Data Processing PipelineabstractWith the interdependencies that exist between data in a scientific processing pipeline, the ability to track the provenance of the scientific process through multiple stages is necessary to determining the usability of the resulting data product. In this paper, we study the capture of provenance from an existing NASA instrument ingest pipeline. Since instrumenting the scientific code for a production system is not feasible, we show how provenance events can be scavenged from log files to generate detailed provenance graphs. Through extensions to the Karma provenance system, which have been implemented on a test instance of the AMSR-E production data pipeline, we determine that when the volume of provenance information is high, provenance graph visualizations provide a good tool for monitoring the ingest pipeline and identifying processing differences in ways not seen before. Two novel uses of provenance that we present in this paper are comparisons between processing runs and forward provenance for viewing downstream dependencies. Scott Jensen, Beth Plale, Mehmet S. Aktas, Peng Chen 0015, Helen Conover |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Hierarchical MapReduce Programming Model and Scheduling AlgorithmsabstractWe present a Hierarchical MapReduce framework that gathers computation resources from different clusters and runs MapReduce jobs across them. The applications implemented in this framework adopt the Map-Reduce-Global Reduce model where computations are expressed as three functions: Map, Reduce, and Global Reduce. Two scheduling algorithms are introduced: Compute Capacity Aware Scheduling for compute-intensive jobs and Data Location Aware Scheduling for data-intensive jobs. Experimental evaluations using a molecule binding prediction tool, Auto Dock, and grep demonstrate promising results for our framework. Beth Plale |
CCGRID | 2 |
| 2012 | Provenance analysis: Towards quality provenanceabstractData provenance, a key piece of metadata that describes the lifecycle of a data product, is crucial in aiding scientists to better understand and facilitate reproducibility and reuse of scientific results. Provenance collection systems often capture provenance on the fly and the protocol between application and provenance tool may not be reliable. As a result, data provenance can become ambiguous or simply inaccurate. In this paper, we identify likely quality issues in data provenance. We also establish crucial quality dimensions that are especially critical for the evaluation of provenance quality. We analyze synthetic and real-world provenance based on these quality dimensions and summarize our contributions to provenance quality. You-Wei Cheah, Beth Plale |
eScience | 2 |
| 2012 | Temporal representation for scientific data provenanceabstractProvenance of digital scientific data is an important piece of the metadata of a data object. It can however grow voluminous quickly because the granularity level of capture can be high. It can also be quite feature rich. We propose a representation of the provenance data based on logical time that reduces the feature space. Creating time and frequency domain representations of the provenance, we apply clustering, classification and association rule mining to the abstract representations to determine the usefulness of the temporal representation. We evaluate the temporal representation using an existing 10 GB database of provenance captured from a range of scientific workflows. Peng Chen 0015, Beth Plale, Mehmet S. Aktas |
eScience | 2 |
| 2012 | Generalized representation and mapping for social-ecological data: Freeing data from the databaseabstractScientific discovery increasingly requires collaboration between scientific sub-domains that often have different representations for their data. To bridge gaps between varying domain representations, researchers are developing metadata and semantic representations meaningful to broader communities. Through exploiting these representations we propose a logical model and architecture by which cross-domain researchers can more easily discover, use, and eventually archive, data. In this paper we present an architecture, intermediate data model, and methodology for mapping diverse social-ecological data sources stored in relational databases to a common representation, and for classifying textual data using machine learning. The results are visualized through client views that are built against the general logical model, and applied against a longitudinal database from social-ecological research. Scott Jensen, Beth Plale, Xiaozhong Liu 0001, David B. Leake, Julie England |
eScience | 2 |
| 2012 | Visualization of network data provenanceabstractVisualization facilitates the understanding of scientific data both through exploration and explanation of the visualized data. Provenance also contributes to the understanding of data by containing the contributing factors behind a result. The visualization of provenance, although supported in existing workflow management systems, generally focuses on small (medium) sized provenance data, lacking techniques to deal with big data with high complexity. This paper discusses visualization techniques developed for exploration and explanation of provenance, including layout algorithm, visual style, graph abstraction techniques, and graph matching algorithm, to deal with the high complexity. We demonstrate through application to two extensively analyzed case studies that involved provenance capture and use over three year projects, the first involving provenance of a satellite imagery ingest processing pipeline and the other of provenance in a large-scale computer network testbed. Peng Chen 0015, Beth Plale, You-Wei Cheah, Devarshi Ghoshal, Scott Jensen |
HiPC | 2 |
| 2012 | Sigiri: uniform resource abstraction for grids and cloudsabstractSUMMARY With the maturation of grid computing facilities and recent explosion of cloud computing data centers, midscale computational science has more options than ever before to satisfy computational needs. But heterogeneity brings complexity. We propose a simple abstraction for interaction with heterogeneous resource managers spanning grid and cloud computing and on features that make the tool useful for the midscale physical or natural scientist. Key strengths of the abstraction are its support for multiple standard job specification languages, preservation of direct user interaction with the service, removing the delay that can come through layers of services, and the predictable behavior under heavy loads. Copyright © 2012 John Wiley & Sons, Ltd. Eran Chinthaka Withana, Beth Plale |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | The Open Provenance Model core specification (v1.1)
Luc Moreau 0001, Ben Clifford, Juliana Freire, Joe Futrelle, Yolanda Gil, Paul Groth, Natalia Kwasnikowska, Simon Miles, Paolo Missier, James D. Myers, Beth Plale, Yogesh L. Simmhan, Eric G. Stephan, Jan Van den Bussche |
Future Gener. Comput. Syst. | 11 |
| 2010 | Streamflow Programming Model for Data Streaming in Scientific WorkflowsabstractGeo-sciences involve large-scale parallel models, high resolution real time data from highly asynchronous and heterogeneous sensor networks and instruments, and complex analysis and visualization tools. Scientific workflows are an accepted approach to executing sequences of tasks on scientists' behalf during scientific investigation. Many geo-science workflows have the need to interact with sensors that produce large continuous streams of data, but programming models provided by scientific workflows are not equipped to handle continuous data streams. This paper proposes a framework that utilizes scientific workflow infrastructure and the benefits of complex event processing to compensate for the impedance mismatch between scientific workflows and continuous data streams. Further we propose and formalize new workflow semantics that would allow the users to not only incorporate stream in scientific workflow, but also make use of the functionalities provided by the complex event processing systems effective within the scientific workflows. Chathura Herath, Beth Plale |
CCGRID | 2 |
| 2010 | WORKEM: Representing and Emulating Distributed Scientific Workflow Execution StateabstractScientific workflows have become an integral part of cyberinfrastructure as their computational complexity and data sizes have grown. However, the complexity of the distributed infrastructure makes design of new workflows, determining the right management policies, debugging, testing or reproduction of errors challenging. Today, workflow engines manage the dependencies between tasks of workflows and there are tools available to wrap scientific codes. There is a need for a customizable, isolated and manageable testing container for design, evaluation and deployment of distributed workflows. To build such an environment, we need to be able to model and represent, capture and possibly reuse the execution flows within each task of a workflow that accurately captures the execution behavior. In this paper, we present the design and implementation of WORKEM, an extensible framework that can be used to represent and emulate workflow execution state. We also detail the use of the framework in two specific case studies (a) design and testing of an orchestration system (b) generation of a provenance database. Our evaluation shows that the framework has minimal overheads and can be scaled to run hundreds of workflows in short durations of time and with a high amount of parallelism. Lavanya Ramakrishnan, Dennis Gannon, Beth Plale |
CCGRID | 3 |
| 2010 | Usage Patterns to Provision for Scientific Experimentation in CloudsabstractDriven by the need to provision resources on demand, scientists are turning to commercial and research test-bed Cloud computing resources to run their scientific experiments. Job scheduling on cloud computing resources, unlike earlier platforms, is a balance between throughput and cost of executions. Within this context, we posit that usage patterns can improve the job execution, because these patterns allow a system to plan, stage and optimize scheduling decisions. This paper introduces a novel approach to utilization of user patterns drawn from knowledge-based techniques, to improve execution across a series of active workflows and jobs in cloud computing environments. Using empirical analysis we establish the accuracy of our prediction approach for two different workloads and demonstrate how this knowledge can be used to improve job executions. Eran Chinthaka Withana, Beth Plale |
CloudCom | 2 |
| 2010 | Trading Consistency for Scalability in Scientific MetadataabstractLong-term repositories that are able to represent the detailed descriptive metadata of scientific data have been recognized as key to both data reuse and preservation of the initial investment in generating the data. Detailed metadata captured during scientific investigation not only enables the efficient discovery of relevant data sets but also is a source for exploring ongoing activity. In XMC Cat metadata catalog, an XML catalog that uses a novel hybrid model to store XML to a relational database, we exploit differences in the temporal utility between browse and search metadata to selectively relax the consistency model used. By ensuring only eventual consistency on parts of the solution, we determine through experimental analysis that the performance and scalability of the catalog can be substantially improved. Scott Jensen, Beth Plale |
eScience | 2 |
| 2010 | Versioning for workflow evolutionabstractScientists working in eScience environments often use workflows to carry out their computations. Since the workflows evolve as the research itself evolves, these workflows can be a tool for tracking the evolution of the research. Scientists can trace their research and associated results through time or even go back in time to a previous stage and fork to a new branch of research. In this paper we introduce the workflow evolution framework (EVF), which is demonstrated through implementation in the Trident workflow workbench. The primary contribution of the EVF is efficient management of knowledge associated with workflow evolution. Since we believe evolution can be used for workflow attribution, our framework will motivate researchers to share their workflows and get the credit for their contributions. Eran Chinthaka Withana, Beth Plale, Roger S. Barga, Nelson Araujo |
HPDC | 2 |
| 2010 | Implementation, performance, and science results from a 30.7 TFLOPS IBM BladeCenter clusterabstractAbstract This paper describes Indiana University's implementation, performance testing, and use of a large high performance computing system. IU's Big Red, a 20.48 TFLOPS IBM e1350 BladeCenter cluster, appeared in the 27th Top500 list as the 23rd fastest supercomputer in the world in June 2006. In spring 2007, this computer was upgraded to 30.72 TFLOPS. The e1350 BladeCenter architecture, including two internal networks accessible to users and user applications and two networks used exclusively for system management, has enabled the system to provide good scalability on many important applications while being well manageable. Implementing a system based on the JS21 Blade and PowerPC 970MP processor within the US TeraGrid presented certain challenges, given that Intel‐compatible processors dominate the TeraGrid. However, the particular characteristics of the PowerPC have enabled it to be highly popular among certain application communities, particularly users of molecular dynamics and weather forecasting codes. A critical aspect of Big Red's implementation has been a focus on Science Gateways, which provide graphical interfaces to systems supporting end‐to‐end scientific workflows. Several Science Gateways have been implemented that access Big Red as a computational resource—some via the TeraGrid, some not affiliated with the TeraGrid. In summary, Big Red has been successfully integrated with the TeraGrid, and is used by many researchers locally at IU via grids and Science Gateways. It has been a success in terms of enabling scientific discoveries at IU and, via the TeraGrid, across the US. Copyright © 2009 John Wiley & Sons, Ltd. Craig A. Stewart, Matthew R. Link, D. Scott McCaulay, Greg Rodgers, George W. Turner, David Y. Hancock, Faisal Saied, Marlon E. Pierce, Ross Aiken, Matthias S. Müller, Matthias Jurenz, Matthias Lieber, Jenett Tillotson, Beth Plale |
Concurr. Comput. Pract. Exp. | 15 |
| 2009 | Application of Management Frameworks to Manage Workflow-Based Systems: A Case Study on a Large Scale E-science ProjectabstractManagement architectures are well discussed in the literature, but their application in real life settings has not been as well covered. Automatic management of a system involves many more complexities than closing the control-loop by reacting to sensor data and executing corrective actions. In this paper, we discuss those complexities and propose solutions to those problems on top of Hasthi management framework, where Hasthi is a robust, scalable, and distributed management framework that enables users to manage a system by enforcing management logic authored by users themselves. Furthermore, we present in detail a real life case study, which uses Hasthi to manage a large, SOA based, e-science cyberinfrastructure. Srinath Perera, Suresh Marru, Thilina Gunarathne, Dennis Gannon, Beth Plale |
ICWS | 5 |
| 2008 | Provenance Collection in an Industry Biochemical Discovery CyberinfrastructureabstractWorkflows are an accepted approach for constructing computational scientific experiments. Provenance capture during workflow execution captures the creation history of datasets. This record is essential to the long-term preservation and reuse of the data, and to making determinations of its quality. We are applying provenance collection to the open source life science grid (LSG) using the Karma tool, and extending the information with semantic information using S-OGSA. The project raises interesting challenges in instrumentation, annotation, and visualization of provenance data. Girish Subramanian, Sribabu Doddapaneni, Beth Plale |
eScience | 4 |
| 2008 | Schema-Independent and Schema-Friendly Scientific Metadata ManagementabstractComputational science is creating a deluge of data, and the automated capture and cataloging of detailed descriptive metadata has been recognized as necessary to enable reuse of this data. Scientific communities in varied disciplines have developed detailed XML metadata schemas to describe data products. Our research has identified characteristics of scientific schemas that can be exploited to efficiently capture and search this metadata based on the schemas specific to each community, but using an easily adaptable framework. Scott Jensen, Beth Plale |
eScience | 2 |
| 2008 | Riding the Geoscience Cyberinfrastructure Wave of Data: Real Time Data Use in Education WorkshopabstractThis workshop brings together scientists, technologists, and educators in a discussion of how data rich geoscience cyberinfrastructure frameworks can be more effectively deployed in high school and early undergraduate settings. Beth Plale |
eScience | 1 |
| 2008 | Special Issue: The First Provenance ChallengeabstractAbstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd. Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009 |
Concurr. Comput. Pract. Exp. | 34 |
| 2008 | Query capabilities of the Karma provenance frameworkabstractAbstract Provenance metadata in e‐Science captures the derivation history of data products generated from scientific workflows. Provenance forms a glue linking workflow execution with associated data products, and finds use in determining the quality of derived data, tracking resource usage, and for verifying and validating scientific experiments. In this article, we discuss the scope of provenance collected in the Karma provenance framework used in the LEAD Cyberinfrastructure project, distinguishing provenance metadata from generic annotations. We further describe our approaches to querying for different forms of provenance in Karma in the context of queries in the first provenance challenge. We use an incremental, building‐block method to construct provenance queries based on the fundamental querying capabilities provided by the Karma service centered on the provenance data model. This has the advantage of keeping the Karma service generic and simple, and yet supports a wide range of queries. Karma successfully answers all but one challenge query. Copyright © 2007 John Wiley & Sons, Ltd. Yogesh L. Simmhan, Beth Plale, Dennis Gannon |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | End-to-End Trustworthy Data Access in Data-Oriented Scientific ComputingabstractData-driven computational science on community computational resources is frequently of a magnitude and scale that it requires that computations be done remotely, generating resulting data collections that are too large to be shipped back to a user's workstation. Service-oriented middleware is well equipped to carry out actions on behalf of a user, but SOA middleware does not address user trust in the privacy of their actions and security of their data. In this paper we develop a model that represents the trust relationship between the users and their remote resources in the grid system. We show how one can construct a trusted relationship from the model, with an emphasis on the importance of context to a specific trust relationship. Sangmi Lee Pallickara, Beth Plale, Dennis Gannon |
CCGRID | 2 |
| 2006 | Calder Query Grid Service: Insights and Experimental EvaluationabstractWe have architected and evaluated a new kind of data resource, one that is composed of a logical collection of ephemeral data streams that could be viewed as a collection of publish-subscribe "channels" over which rich data-access and semantic operations can be performed. This paper contributes new insight to stream processing under the highly asynchronous stream workloads often found in data-driven scientific applications, and presents insights gained through porting a distributed stream processing system to a grid services framework. Experimental results reveal limits on stream processing rates that are directly tied to differences in stream rates. Nithya N. Vijayakumar, Ying Liu 0043, Beth Plale |
CCGRID | 3 |
| 2006 | A Framework for Collecting Provenance in Data-Centric Scientific WorkflowsabstractThe increasing ability for the Earth sciences to sense the world around us is resulting in a growing need for data-driven applications that are under the control of data-centric workflows composed of grid- and Web-services. The focus of our work is on provenance collection/or these workflows, necessary to validate the workflow and to determine quality of generated data products. The challenge we address is to record uniform and usable provenance metadata that meets the domain needs while minimizing the modification burden on the service authors and the performance overhead on the workflow engine and the services. The framework, based on a loosely-coupled publish-subscribe architecture for propagating provenance activities, satisfies the needs of detailed provenance collection while a performance evaluation of a prototype finds a minimal performance overhead (in the range of 1% for an eight service workflow using 271 data products) Yogesh L. Simmhan, Beth Plale, Dennis Gannon |
ICWS | 2 |
| 2006 | Dynamic Filtering and Mining Triggers in Mesoscale Meteorology ForecastingabstractAbstract — Mesoscale meteorology forecasting as a data driven application is capable of reacting to events in real-time. We explore a framework for dynamic filtering and mining of data products to generate timely triggers for invoking forecasting applications. In this paper, we present our framework, which couples the Calder stream processing system developed at Indiana University for filter processing and trigger generation, and data mining algorithms developed as part of the ADaM data mining tool kit developed at ITSC, UAH, which detect events for trigger generation. Nithya N. Vijayakumar, Beth Plale, Rahul Ramachandran, Xiang Li 0043 |
IGARSS | 2 |
| 2006 | Bandwidth challenge - All in a day's work: advancing data-intensive research with the data capacitorabstractIndiana University provides powerful compute, storage, and network resources to a diverse local and national research community every day. IU's facilities have been used to support data-intensive applications ranging from digital humanities to computational biology.For this year's bandwidth challenge, several IU researchers will conduct experiments from the exhibit floor utilizing the resources that University Information Technology Services currently provides.Using IU's newly constructed 535 TB Data Capacitor and an additional component installed on the exhibit floor, we will use Lustre across the wide area network to simultaneously facilitate dynamic weather modeling, protein analysis, instrument data capture, and the production, storage, and analysis of simulation data. Stephen C. Simms, Matt Davy, Bret Hammond, Matthew R. Link, Craig A. Stewart, Randall Bramley, Beth Plale, Dennis Gannon, Mu-Hyun Baik, Scott Teige, John C. Huffman, Rick McMullen, Doug Balog, Gregory G. Pike |
SC | 7 |
| 2006 | Poster reception - A meta-provenance service to infer context from provenance data of distributed entitiesabstractProvenance management has become an integral part of many large-scale distributed computing systems. Tracking the history of data and its usage has led to better understanding of system requirements as well as user needs. Still, the need for an intelligent service that matches the system requirements with user needs is not satisfied. We propose a meta-provenance service that infers context from the provenance information of distributed entities and uses this contextual information to satisfy user needs. We describe our meta-provenance framework by way of describing its implementation in the Calder system. The Calder streaming system enables dynamic invocation of forecast models in LEAD by using a distributed mesh of data mining agents. The meta-provenance service enables sophisticated mapping of user queries from the LEAD portal down to the set of few data mining agents that execute them. Also our meta-provenance service can work at multiple levels of contextual granularity. Nithya N. Vijayakumar, Beth Plale |
SC | 2 |
| 2006 | Multi-model Based Optimization for Stream Query Processing
Ying Liu 0043, Beth Plale |
SEKE | 2 |
| 2005 | Distributed streaming query planner in Calder systemabstractThe contribution of this work has two folds. First, we extend the current query planners' cost metric space by introducing network bandwidth cost, query deployment cost and query re-using cost; second, we develop a suite of algorithms for re-using existing query fragments under different scenarios. One of the most important reusable queries is called structure-sharable query. Ying Liu 0043, Beth Plale, Nithya N. Vijayakumar |
HPDC | 2 |
| 2005 | Calder: enabling grid access to data streamsabstractThis paper presents an experimental evaluation of a grid-based continuous query solution to access data streams. The results presented in this study are mixed, however. We are migrating to a netCDF-based streaming model to measure system behavior in a setting that more accurately reflects a real use scenario. We are also working on enabling approximate query processing support to Calder to deal with sudden drop offs and changes in stream rates. Nithya N. Vijayakumar, Ying Liu 0043, Beth Plale |
HPDC | 3 |
| 2005 | Service Oriented Architectures for Science Gateways on Grid Systems
Dennis Gannon, Beth Plale, Marcus Christie, Scott Jensen, Gopi Kandaswamy, Suresh Marru, Sangmi Lee Pallickara, Satoshi Shirasuna, Yogesh L. Simmhan, Aleksander Slominski, Yiming Sun 0001 |
ICSOC | 2 |
| 2005 | Building Grid Portal Applications From a Web Service Component ArchitectureabstractThis work describes an approach to building Grid applications based on the premise that users who wish to access and run these applications prefer to do so without becoming experts on Grid technology. We describe an application architecture based on wrapping user applications and application workflows as Web services and Web service resources. These services are visible to the users and to resource providers through a family of Grid portal components that can be used to configure, launch, and monitor complex applications in the scientific language of the end user. The applications in this model are instantiated by an application factory service. The layered design of the architecture makes it possible for an expert to configure an application factory service with a custom user interface client that may be dynamically loaded into the portal. Dennis Gannon, Jay Alameda, Octav Chipara, Marcus Christie, Vinayak Dukle, Matthew Farrellee, Gopi Kandaswamy, Deepti Kodeboyina, Sriram Krishnan, Charles W. Moad, Marlon E. Pierce, Beth Plale, Albert L. Rossi, Yogesh L. Simmhan, Anuraag Sarangi, Aleksander Slominski, Satoshi Shirasuna, Thomas Thomas |
Proc. IEEE | 13 |
| 2004 | Understanding Grid resource information management through a synthetic database benchmark/workloadabstractManagement of Grid resource information is a challenging, and important area considering the potential size of the Grid and wide range of resources that should be represented. Though example Grid Information Servers exist, behavior of these servers across different platforms is less well understood. This paper describes a study we undertook to compare the access language and platform capabilities for three different database platforms, relational, native XML, and LDAP, serving as a Grid information server. Our study measures query response times for a range of queries and highlights sensitivities exhibited by the different platforms to variables such as result set size and collection size. Beth Plale, Craig Jacobs, Scott Jensen, Ying Liu 0043, Charles W. Moad, Rupali Parab, Prajakta Vaidya |
CCGRID | 1 |
| 2004 | Building Grid Applications and Portals: An Approach Based on Components, Web Services and Workflow Tools
Dennis Gannon, Gopi Kandaswamy, Deepti Kodeboyina, Sriram Krishnan, Beth Plale, Aleksander Slominski |
Euro-Par | 6 |
| 2004 | Performance Evaluation of Rate-Based Join Window Sizing for Asynchronous Data Streams
Nithya N. Vijayakumar, Beth Plale |
HPDC | 2 |
| 2003 | Dynamic Querying of Streaming Data with the dQUOB SystemabstractData streaming has established itself as a viable communication abstraction in data-intensive parallel and distributed computations, occurring in applications such as scientific visualization, performance monitoring, and large-scale data transfer. A known problem in large-scale event communication is tailoring the data received at the consumer. It is the general problem of extracting data of interest from a data source, a problem that the database community has successfully addressed with SOL queries, a time tested, user-friendly way for noncomputer scientists to access data. By leveraging the efficiency of query processing provided by relational queries, the dQUOB system provides a conceptual relational data model and SOL query access over streaming data. Queries can be used to extract data, combine streams, and create new streams. The language augments queries with an action to enable more complex data transformations such as Fourier transforms. The dQUOB system has been applied to two large-scale distributed applications: a safety critical autonomous robotics simulation and scientific software visualization for global atmospheric transport modeling. In this paper, we present the dQUOB system and the results of performance evaluation undertaken to assess its applicability in data-intensive wide-area computations, where the benefit of portable data transformation must be evaluated against the cost of continuous query evaluation. Beth Plale, Karsten Schwan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | Leveraging Run Time Knowledge about Event Rates to Improve Memory Utilization in Wide Area Data Stream FilteringabstractThe dQUOB system conceptualization of data streams as database and its SQL interface to data streams is an intuitive way for users to think about their data needs in a large scale application containing hundreds if not thousands of data streams. Experience with dQUOB has shown the need for more aggressive memory management to achieve the scalability we desire. This paper addresses the problem with a two-fold solution. The first one is replacement of the existing first-come first-served scheduling algorithm with an earliest job first algorithm which we demonstrate to yield better average service time. The second one is an introspection algorithm that sets and adapts the sizes of join windows in response to the knowledge acquired at runtime about event rates. In addition to the potential for significant improvements in memory utilization, the algorithm presented here also provides a means by which the user can reason about join window sizes. Wide area measurements demonstrate the adaptive capability required by the introspection technique. Beth Plale |
HPDC | 1 |
| 2001 | Optimizations Enabled by Relational Data Model View to Querying Data StreamsabstractWe postulate that the popularity and efficiency of SQL for querying relational databases makes the language a viable solution to retrieving data from data streams. In response, we have developed a system, dQUOB, that uses SQL queries to extract data from streaming data in real time. The high performance needs of applications such as scientific visualization motivates our search for optimizations to improve query evaluation efficiency. The purpose of this paper is to discuss the unique optimizations we have realized by a database point of view to streaming data and to show that the enhanced conceptual model of viewing data streams as relations has reasonable overhead. Beth Plale, Karsten Schwan |
IPDPS | 1 |
| 2001 | Taking the Step From Meta-Information to Communication Middleware in Computational Data StreamsabstractIt is our belief that network applications relying on globally distributed shared resources will increasingly adopt meta-level descriptions to describe the data streaming in the application at runtime. Our group has developed the notion of computational data streams to describe and act upon such data flows. Our work conceptualizes the data flows as database relations over which useful operations, such as querying, can be performed. This paper shows how one can make the step from a meta-level description of data flows to an actual implementation using CORBA-style event channels and binary I/O for data transport. 1 Beth Plale, Patrick M. Widener, Karsten Schwan |
IPDPS | 1 |
| 2000 | dQUOB: Managing Large Data Flows using Dynamic Embedded QueriesabstractThe dQUOB system satisfies client need for specific information from high-volume data streams. The data streams we speak of are the flow of data existing during large-scale visualizations, video streaming to large numbers of distributed users, and high volume business transactions. We introduce the notion of conceptualizing a data stream as a set of relational database tables so that a scientist can request information with an SQL-like query. Transformation or computation that often needs to be performed on the data en-route can be conceptualized as computation performed on consecutive views of the data, with computation associated with each view. The dQUOB system moves the query code into the data stream as a quoblet; as compiled code. The relational database data model has the significant advantage of presenting opportunities for efficient reoptimizations of queries and sets of queries. Using examples from global atmospheric modeling, we illustrate the usefulness of the dQUOB system. We carry the examples through the experiments to establish the viability of the approach for high performance computing with a baseline benchmark. We define a cost-metric of end-to-end latency that can be used to determine realistic cases where optimization should be applied. Finally, we show that end-to-end latency can be controlled through a probability assigned to a query that a query will evaluate to true. Beth Plale, Karsten Schwan |
HPDC | 1 |
| 1999 | Steering Data Streams in Distributed Computational LaboratoriesabstractThis research supports the interactive access to large-scale scientific data by creation of active user interfaces (AUIs). An AUI continuously emits events describing its current information needs, based on which methods may be developed for controlling the potentially immense information streams directed at the interface. More precisely the purposes of stream control are twofold. First, stream control is performed to deal with heterogeneity in underlying systems, where low end displays may receive only small portions of the data shown at high end displays. Second, stream control is used to achieve scalability with respect to the size and complexity of data streams directed at a user interface, by filtering the data stream and by offloading certain computations from the AUI to the information generators or to information routing sites, by dynamically migrating such computations to appropriate locations, and by adapting these computations in order to effect tradeoffs in the amount of data moved across network links vs. the computations required. Carsten Isert, Davis King 0001, Karsten Schwan, Beth Plale, Greg Eisenhauer |
HPDC | 4 |
| 1999 | Run-time Detection in Parallel and Distributed Systems: Application to Safety-Critical SystemsabstractThere is growing interest in run-time detection as parallel and distributed systems grow larger and more complex. This work targets run-time analysis of complex, interactive scientific applications for purposes of attaining scalability improvements with respect to the amount and complexity of the data transmitted, transformed, and shared among different application components. Such improvements are derived from using database techniques to manipulate data streams. Namely, by imposing a relational model on the data streams, constraints on the stream may be expressed as database queries evaluated against the data events comprising the stream. The application in the paper is to a safety-critical system. The paper also presents a tool, dQUOB, Dynamic QUery OBjects, which: (1) offers the means for dynamic creation of queries and for their application to large data streams; (2) permits implementation and runtime use of multiple "query optimization" techniques; and (3) supports dynamic reoptimization of queries based on streams' dynamic behavior. Beth Plale, Karsten Schwan |
ICDCS | 1 |
| 1998 | DataExchange: High Performance Communications in Distributed Laboratories
Greg Eisenhauer, Beth Plale, Karsten Schwan |
Parallel Comput. | 2 |