David Chiu 0001

dblp:64/411 · DBLP profile ↗
← Back
32ranked-venue papers
9as first author
3since 2021 · last 2023
0000-0002-1210-8268ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 4 first-author · 3 since 2021Systems, architecture and hardware · 9 · 3 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1
YearPublicationVenuePosition
2023 Workload-Aware Cache Management of Bitmap Indices
abstract
Big-data management systems must handle multiple concurrent queries over multi-dimensional data sets. To achieve high throughput, such systems could implement various techniques to avoid redundant computations and data fetches. One such approach is to cache a subset of the query results and reuse these results to (partially) fulfill future query requests. This approach can be quite effective for query-at-a-time processing. However, we suspect that even greater performance is being left on the table if queries are only optimized in isolation, and that higher throughput can be extracted through a systematic examination of the relationships between queries in a given workload.
Julia Kaeppel, Jason Sawin, David Chiu 0001
BDCAT3
2021 Caching Support for Range Query Processing on Bitmap Indices
abstract
Bitmaps are commonly used for indexing read-mostly data sets. The range of an attribute is split into bins, where its values are placed: bij = 1 denotes the value of the ith tuple is in the jth bin, and bij = 0 otherwise. A number of query types can be decomposed into the systematic application of boolean operators over sets of bins. However, when bitmaps are high-dimensional, the overall query-processing performance can deteriorate due to the increased number of bins that participate per query.
Sarah McClain, Manya Mutschler-Aldine, Colin Monaghan, David Chiu 0001, Jason Sawin, Patrick Jarvis
SSDBM4
2021 Exploring Means to Enhance the Efficiency of GPU Bitmap Index Query Processing
abstract
Abstract Once exotic, computational accelerators are now commonly available in many computing systems. Graphics processing units (GPUs) are perhaps the most frequently encountered computational accelerators. Recent work has shown that GPUs are beneficial when analyzing massive data sets. Specifically related to this study, it has been demonstrated that GPUs can significantly reduce the query processing time of database bitmap index queries. Bitmap indices are typically used for large, read-only data sets and are often compressed using some form of hybrid run-length compression. In this paper, we present three GPU algorithm enhancement strategies for executing queries of bitmap indices compressed using word aligned hybrid compression: (1) data structure reuse (2) metadata creation with various type alignment and (3) a preallocated memory pool. The data structure reuse greatly reduces the number of costly memory system calls. The use of metadata exploits the immutable nature of bitmaps to pre-calculate and store necessary intermediate processing results. This metadata reduces the number of required query-time processing steps. Preallocating a memory pool can reduce or entirely remove the overhead of memory operations during query processing. Our empirical study showed that performing a combination of these strategies can achieve 32.4 $$\times$$ × to 98.7 $$\times$$ × speedup over the current state-of-the-art implementation. Our study also showed that by using our enhancements, a common gaming GPU can achieve a $$15.0\times$$ 15.0 × speedup over a more expensive high-end CPU.
Brandon Tran, Brennan Schaffner, Joe Myre, Jason Sawin, David Chiu 0001
Data Sci. Eng.5
2020 Increasing the Efficiency of GPU Bitmap Index Query Processing
Brandon Tran, Brennan Schaffner, Jason Sawin, Joe Myre, David Chiu 0001
DASFAA (3)5
2019 A Transactional Framework for Broadening Access to Geo-Diversification
abstract
This paper establishes a transactional energy market in which any data center can participate. Like a stock-market exchange, we propose a framework in which data centers can trade energy usage (in the form of jobs) for monetary value. The proposed framework allows each data center to monitor multiple parameters, including the current energy prices, budgets, and job execution states. These parameters inform the construction of models to help participating data centers optimize various cost, profit, and job-performance objectives to manage the risks of market participation. In our feasibility study, market participants mutually benefit by increasing general revenue through either a reduction of energy costs and/or through the successful completion of more jobs. Using real energy-pricing data, in our simulated experiment of 100 participating data centers, we observe a significant cost reduction leading to an average increase of 17.8% profit margins.
Jared Polonitza, David Chiu 0001
CLOUD2
2019 GPU Acceleration of Range Queries over Large Data Sets
abstract
Data management systems commonly use bitmap indices to increase the efficiency of querying scientific data. Bitmaps are usually highly compressible and can be queried directly using fast hardware-supported bitwise logical operations. The processing of bitmap queries is inherently parallel in structure, which suggests they could benefit from concurrent computer systems. In particular, bitmap-range queries offer a highly parallel computational problem, and the hardware features of graphics processing units (GPUs) offer an alluring platform for accelerating their execution. In this paper, we present three GPU algorithms and one CPU based algorithm for the parallel execution of bitmap-range queries. We show that in 95% of our tests, using real and synthetic data, the GPU algorithms greatly outperform the parallel CPU algorithm. For these tests, the GPU algorithms provide up to 87.7× speedup and an average speedup of 30.22× over the parallel CPU algorithm. In addition to enhancing performance, augmenting traditional bitmap query systems with GPUs to offload bitmap query processing allows the CPU to process other requests.
Mitchell Nelson, Zachary Sorenson, Joe Myre, Jason Sawin, David Chiu 0001
BDCAT5
2018 Fault-Tolerant Query Execution over Distributed Bitmap Indices
abstract
Advances in storage software and filesystems have proliferated a vast array of easy-to-use distributed storage services, removing the barrier for a growing number of organizations to geo-distribute large data sets. While leaving data in their distributed environments is convenient for data collection, various types of processing (that might use multiple data sources) are precluded due to the prohibitive costs of data movement. Users are therefore burdened with finding creative ways of performing data analysis, often requiring expert knowledge in multiple domains. This paper reports on the design and implementation of a query engine that enables high-level queries over distributed data sets. Our system generates bitmap indices at multiple geo-distributed data sources in order to approximate large amounts of raw data values. The bitmaps are replicated for fault-tolerance and performance. Upon accepting a high-level (SQL-like) query, our system generates a query plan, resolves dependencies, and schedules for its execution over the distributed system. The system has been tested rigorously, and experimental results show that most overheads (i.e., query planning, node spawning, etc.) are negligible. Our testing also shows that our system is capable of delivering query results in the face of node failures, with no observable impact on query execution for up to 20% of the system failing. The system also provides a framework that is easily extendible for future research on the interplay between distributed systems and bitmap indices.
Sam Burdick, Jahrme Risner, David Chiu 0001, Jason Sawin
BDCAT3
2017 Improving the Qerying Efficiency of the PLWAH Bitmap Algorithm
abstract
Bitmap indices are commonly used for accessing large, read-only data. A bitmap is a simplified model of the underlying data in secondary storage. Its coarse representation enables the use of fast CPU operations to answer common database queries. Additionally, bitmaps are very compressible. Several known compression algorithms allow the compressed form of the bitmap to be queried directly, and one of which is Position List Word-Aligned Hybrid (PLWAH). PLWAH is modified hybrid run-length encoding scheme that can achieve better compression than traditional schemes such as Word-Aligned Hybrid (WAH). This improved compression introduces an increased query processing cost, of which we address in this paper. We present a technique that uses metadata to allow PLWAH's query algorithm to exploit logical short-circuiting opportunities, reducing the cost of certain queries. In our empirical study, we found that our approach achieved an average speedup of 1.41x over PLWAH for real scientific data sets. For specific queries, our approach realized speedups as high as 8000x.
Benjamin Taufen, Jason Sawin, David Chiu 0001
IDEAS3
2016 Hadoop in Flight: Migrating Live MapReduce Jobs for Power-Shifting Data Centers
abstract
Renewable energy sources such as wind and solar are unpredictable for power utilities, which must produce exactly as much power as is needed at any given time. To help manage the demand, some utilities have begun deploying real-time energy prices to their customers. Data centers, which often run Hadoop jobs on thousands of machines, have become some of the utilities' largest consumers. In fact, recent studies have shown that, when processing at full capacity, data centers can require as much power as a mid-sized U. S. city. By implementing a method in which data centers can offload their work to locations on different power grids, they can take advantage of the lower-priced energy and thereby minimize operational costs. To this end, we have designed and implemented a new mechanism directly within the Hadoop 2 codebase that allows users to pause, migrate, and resume a job at arbitrary points of execution. We have evaluated this scheme using popular applications and show that energy can be delayed and shifted to a different location with reasonable overheads. Our experiments justify the migration use-case, showing that it saves both energy and time over either restarting the job remotely or allowing it to complete locally.
Chili Johnson, David Chiu 0001
CLOUD2
2016 A Two-Phase MapReduce Algorithm for Scalable Preference Queries over High-Dimensional Data
abstract
Preference (top-k) queries play a key role in modern data analytics tasks. Top-k techniques rely on ranking functions in order to determine an overall score for each of the objects across all the relevant attributes being examined. This ranking function is provided by the user at query time, or generated for a particular user by a personalized search engine which prevents the pre-computation of the global scores. Executing this type of queries is particularly challenging for high-dimensional data. Recently, bit-sliced indices (BSI) were proposed to answer these high-dimensional preference queries efficiently in a centralized environment.
Gheorghi Guzun, Guadalupe Canahuate, David Chiu 0001
IDEAS3
2015 Topic Model for Identifying Suicidal Ideation in Chinese Microblog
Xiaolei Huang 0002, Tianli Liu, David Chiu 0001, Tingshao Zhu
PACLIC4
2014 A tunable compression framework for bitmap indices
abstract
Bitmap indices are widely used for large read-only repositories in data warehouses and scientific databases. Their binary representation allows for the use of bitwise operations and specialized run-length compression techniques. Due to a trade-off between compression and query efficiency, bitmap compression schemes are aligned using a fixed encoding length size (typically the word length) to avoid explicit decompression during query time. In general, smaller encoding lengths provide better compression, but require more decoding during query execution. However, when the difference in size is considerable, it is possible for smaller encodings to also provide better execution time. We posit that a tailored encoding length for each bit vector will provide better performance than a one-size-fits-all approach. We present a framework that optimizes compression and query efficiency by allowing bitmaps to be compressed using variable encoding lengths while still maintaining alignment to avoid explicit decompression. Efficient algorithms are introduced to process queries over bitmaps compressed using different encoding lengths. An input parameter controls the aggressiveness of the compression providing the user with the ability to tune the tradeoff between space and query time. Our empirical study shows this approach achieves significant improvements in terms of both query time and compression ratio for synthetic and real data sets. Compared to 32-bit WAH, VAL-WAH produces up to 1.8× smaller bitmaps and achieves query times that are 30% faster.
Gheorghi Guzun, Guadalupe Canahuate, David Chiu 0001, Jason Sawin
ICDE3
2014 Lightweight online power monitoring and control for mobile applications
abstract
Limited battery power has long been a challenge for mobile applications. As a result, the work in power monitoring and management has attracted great interests. In this paper, we propose a model to estimate power consumption of mobile applications at run-time, based on application-specific per-action power profiling. In addition, we have developed on-line optimization techniques which help maximize users' experience while conserving power. Our power model is lightweight and flexible, in that it can be used by any mobile applications as a plugin, and it can support user-defined optimization mechanisms. This approach has been evaluated using a case study, a mobile application for field studies, and the experimental results show that our model accurately captures power consumption of the application, and the model can be used to optimize the power consumption based on users' needs.
Xinghui Zhao, David Chiu 0001
ICPADS3
2014 Design and evaluation of a Twitter hashtag recommendation system
abstract
Twitter has evolved into a powerful communication and information sharing tool used by millions of people around the world to post what is happening now. A hashtag, a keyword prefixed with a hash symbol (#), is a feature in Twitter to organize tweets and facilitate effective search among a massive volume of data. In this paper, we propose an automatic hashtag recommendation system that helps users find new hashtags related to their interests.
Eriko Otsuka, Scott A. Wallace, David Chiu 0001
IDEAS3
2014 Optimizing query execution for variable-aligned length compression of bitmap indices
abstract
Indexing is a fundamental mechanism for efficient data access. Recently, we proposed the Variable-Aligned Length (VAL) bitmap index encoding framework, which generalizes the commonly used word-aligned compression techniques. VAL presented a variable-aligned compression framework, which allows columns of a bitmap to be compressed using different encoding lengths. This flexibility creates a tunable compression that balances the trade-off between space and query processing time. The variable format of VAL presents several unique opportunities for query optimization.
Ryan Slechta, Jason Sawin, Ben McCamish, David Chiu 0001, Guadalupe Canahuate
IDEAS4
2014 Poster: A power-aware mobile app for field scientists
abstract
In this poster, we design and implement a mobile application and back-end management system to help field scientists manage data collection, improve real-time communication, and optimize power consumption during a scientific field study.
Xinghui Zhao, David Chiu 0001
MobiSys3
2013 Dynamic bitmap index recompression through workload-based optimizations
abstract
Many large-scale read-only databases and data warehouses use bitmap indices in an effort to speed up data analysis. These indices have the dual properties of compressibility and being able to leverage fast bit-wise operations for query processing. Numerous hybrid run-length encoding compression schemes have been proposed that greatly compress the index and enable querying without the need to decompress. Typically, these schemes align their compression with the computer architecture's word size to further accelerate queries.
Fredton Doan, David Chiu 0001, Brasil Perez Lukes, Jason Sawin, Gheorghi Guzun, Guadalupe Canahuate
IDEAS2
2013 Integrating Online Compression to Accelerate Large-Scale Data Analytics Applications
abstract
Compute cycles in high performance systems are increasing at a much faster pace than both storage and wide-area bandwidths. To continue improving the performance of large-scale data analytics applications, compression has therefore become promising approach. In this context, this paper makes the following contributions. First, we develop a new compression methodology, which exploits the similarities between spatial and/or temporal neighbors in a popular climate simulation dataset and enables high compression ratios and low decompression costs. Second, we develop a framework that can be used to incorporate a variety of compression and decompression algorithms. This framework also supports a simple API to allow integration with an existing application or data processing middleware. Once a compression algorithm is implemented, this framework automatically mechanizes multi-threaded retrieval, multi-threaded data decompression, and the use of informed prefetching and caching. By integrating this framework with a data-intensive middleware, we have applied our compression methodology and framework to three applications over two datasets, including the Global Cloud-Resolving Model (GCRM) climate dataset. We obtained an average compression ratio of 51.68%, and up to 53.27% improvement in execution time of data analysis applications by amortizing I/O time by moving compressed data.
Tekin Bicer, Jian Yin 0002, David Chiu 0001, Gagan Agrawal, Karen Schuchardt
IPDPS3
2013 Cost and Accuracy Aware Scientific Workflow Composition for Service-Oriented Environments
abstract
Large-scale scientific data analysis projects have catalyzed service-based workflow management systems. We present an approach for integrating user preferences on completion time and workflow accuracy in a workflow composition system. The relationship between workflow execution time and the accuracy of results is exploited by our workflow system. Specifically, our system is equipped with a way for users to define cost models on service completion time and error propagation (prevalent in many scientific and data analysis applications). Together with these models and an ontology for describing web service and data dependences, our system plans service-based workflows to answer high-level queries. Our system was evaluated under a real service-based environment against user constraints on time, accuracy, and network bandwidth variations. In the worst case in our experiments, we observed an average deviation of 14.3 percent below the desired time constraints, which suggests that our system is time-conservative. Within varying network bandwidth environments, we can also meet time constraints through sampling, and only a 12.4 percent deviation below time expectations are observed on average. We further show that, though negotiating with services' error models, our system is capable of planning data reduction measures (e.g., sampling) directly within workflow plans to achieve the desired accuracy.
David Chiu 0001, Gagan Agrawal
IEEE Trans. Serv. Comput.1
2012 Time and Cost Sensitive Data-Intensive Computing on Hybrid Clouds
abstract
Purpose-built clusters permeate many of today's organizations, providing both large-scale data storage and computing. Within local clusters, competition for resources complicates applications with deadlines. However, given the emergence of the cloud's pay-as-you-go model, users are increasingly storing portions of their data remotely and allocating compute nodes on-demand to meet deadlines. This scenario gives rise to a hybrid cloud, where data stored across local and cloud resources may be processed over both environments. While a hybrid execution environment may be used to meet time constraints, users must now attend to the costs associated with data storage, data transfer, and node allocation time on the cloud. In this paper, we describe a modeling-driven resource allocation framework to support both time and cost sensitive execution for data-intensive applications executed in a hybrid cloud setting. We evaluate our framework using two data-intensive applications and a number of time and cost constraints. Our experimental results show that our system is capable of meeting execution deadlines within a 3.6% margin of error. Similarly, cost constraints are met within a 1.2% margin of error, while minimizing the application's execution time.
Tekin Bicer, David Chiu 0001, Gagan Agrawal
CCGRID2
2012 Compiler and runtime support for enabling reduction computations on heterogeneous systems
abstract
SUMMARY A trend that has materialized, and has given rise to much attention, is of the increasingly heterogeneous computing platforms. Presently, it has become very common for a desktop or a notebook computer to come equipped with both a multi‐core CPU and a graphics processing unit (GPU). Capitalizing on the maximum computational power of such architectures (i.e., by simultaneously exploiting both the multi‐core CPU and the GPU), starting from a high‐level API, is a critical challenge. We believe that it would be highly desirable to support a simple way for programmers to realize the full potential of today's heterogeneous machines. This paper describes a compiler and runtime framework that can map a class of applications, namely those characterized bygeneralized reductions, to a system with a multi‐core CPU and GPU. Starting with simple C functions with added annotations, we automatically generate the middleware API code for the multi‐core, as well as CUDA code to exploit the GPU simultaneously. The runtime system provides efficient schemes for dynamically partitioning the work between CPU cores and the GPU. Our experimental results from two applications, for example, k‐means clustering and principal component analysis, show that, through effectively harnessing the heterogeneous architecture, we can achieve significantly higher performance compared with using only the GPU or the multi‐core CPU. In k‐means clustering, the heterogeneous version with eight CPU cores and a GPU achieved a speedup of about 32.09x relative to one‐thread CPU. When compared with the faster of CPU‐only and GPU‐only executions, we were able to achieve a performance gain of about 60%. In principal component analysis, the heterogeneous version attained a speedup of 10.4x relative to the one‐thread CPU version. When compared with the faster of CPU‐only and GPU‐only versions, the heterogeneous version achieved a performance gain of about 63.8%. Copyright © 2011 John Wiley & Sons, Ltd.
Vignesh T. Ravi, Wenjing Ma, David Chiu 0001, Gagan Agrawal
Concurr. Comput. Pract. Exp.3
2011 Evaluating and Optimizing Indexing Schemes for a Cloud-Based Elastic Key-Value Store
abstract
Cloud computing has emerged to provide virtual, pay-as-you-go computing and storage services over the Internet, where the usage cost directly depends on consumption. One compelling feature in Clouds is elasticity, where a user can demand, and be immediately given access to, more (or less) resources based on requirements. However, this feature introduces new challenges in developing application and services. In this paper, we focus on the challenges in data management in Cloud environments, in view of elasticity. Particularly, we consider an elastic key-value store, which is used to cache intermediate results in a service-oriented system, and accelerate future queries by reusing the stored values. Such a key-value store can clearly benefit from the elasticity offered by Clouds, by expanding the cache during query-intensive periods. However, supporting an elastic key-value store involves many challenges, including selecting an appropriate indexing scheme, data migration upon elastic resource provisioning, and optimizations to remove certain overheads in the Cloud. This paper focuses on the design of an elastic key-value store. We consider three ubiquitous methods for indexing: B+-Trees, Extendible Hashing, and Bloom Filters, and we show how these schemes can be modified to exploit elasticity in Clouds. We also evaluate various performance aspects associated with the use of these indexing schemes. Furthermore, we have developed a heuristic to request elastic compute resources for expanding the cache such that instance startup overheads are minimized in our scheme. Our evaluation studies show that the index selection depends on various application and system level parameters that we have identified. And while we confirm that B+-Trees, which pervade many of today's key-value systems, would scale well, we showcases when Extendible Hashing would outperform B+-Trees.
David Chiu 0001, Apeksha Shetty, Gagan Agrawal
CCGRID1
2011 A Framework for Data-Intensive Computing with Cloud Bursting
abstract
For many organizations, one attractive use of cloud resources can be through what is referred to as cloud bursting or the hybrid cloud. These refer to scenarios where an organization acquires and manages in-house resources to meet its base need, but can use additional resources from a cloud provider to maintain an acceptable response time during workload peaks. Cloud bursting has so far been discussed in the context of using additional computing resources from a cloud provider. However, as next generation applications are expected to see orders of magnitude increase in data set sizes, cloud resources can be used to store additional data after local resources are exhausted. In this paper, we consider the challenge of data analysis in a scenario where data is stored across a local cluster and cloud resources. We describe a software framework to enable data-intensive computing with cloud bursting, i.e., using a combination of compute resources from a local cluster and a cloud environment to perform Map-Reduce type processing on a data set that is geographically distributed. Our evaluation with three different applications shows that data-intensive computing with cloud bursting is feasible and scalable. Particularly, as compared to a situation where the data set is stored at one location and processed using resources at that end, the average slowdown of our system (using distributed but the same aggregate number of compute resources), is only 15.55%. Thus, the overheads due to global reduction, remote data retrieval, and potential load imbalance are quite manageable. Our system scales with an average speedup of 81% when the number of compute resources is doubled.
Tekin Bicer, David Chiu 0001, Gagan Agrawal
CLUSTER2
2011 Variable Length Compression for Bitmap Indices
Fabian Corrales, David Chiu 0001, Jason Sawin
DEXA (2)2
2011 An approach towards automatic workflow composition through information retrieval
abstract
Understanding how to design, manage, and execute scientific workflows has become increasingly esoteric. Yet, despite the development of scientific workflow management systems, which have simplified workflow planning to some extent, a means to reduce the complexity of user interaction without forfeiting some robustness has been elusive. We believe that a keyword interface may be highly beneficial to common users in need of information which requires workflow planning and execution. In this paper, we describe a system that can automatically compose a set of relevant workflows, which may or may not have been previously defined by other users, given only a keyword query. We present a way to index data sets and Web services (utilized to compose workflows in our system) on their ontological attributes. This ontology allows us to facilitate an IR-based workflow retrieval model. We conducted a case study in geoinformatics with a set of real geospatial Web services, data, and their metadata annotations. our system was capable of answering six keyword queries with fast search times (2.16ms on average) and relatively high Top-N precision values: 78%, 77.3%, and 76.2% for the Top 3, 5, and 10 retrieved workflows respectively.
David Chiu 0001, Travis Hall, Farhana Kabir, Gagan Agrawal
IDEAS1
2011 Keyword Search Support for Automating Scientific Workflow Composition
David Chiu 0001, Travis Hall, Farhana Kabir, Gagan Agrawal
SSDBM1
2010 Compiler and runtime support for enabling generalized reduction computations on heterogeneous parallel configurations
abstract
A trend that has materialized, and has given rise to much attention, is of the increasingly heterogeneous computing platforms. Presently, it has become very common for a desktop or a notebook computer to come equipped with both a multi-core CPU and a GPU. Capitalizing on the maximum computational power of such architectures (i.e., by simultaneously exploiting both the multi-core CPU and the GPU) starting from a high-level API is a critical challenge. We believe that it would be highly desirable to support a simple way for programmers to realize the full potential of today's heterogeneous machines.
Vignesh T. Ravi, Wenjing Ma, David Chiu 0001, Gagan Agrawal
ICS3
2010 Elastic Cloud Caches for Accelerating Service-Oriented Computations
abstract
Computing as a utility, that is, on-demand access to computing and storage infrastructure, has emerged in the form of the Cloud. In this model of computing, elastic resource allocation, i.e., the ability to scale resource allocation for specific applications, should be optimized to manage cost versus performance. Meanwhile, the wake of the information sharing/mining age is invoking a pervasive sharing of Web services and data sets in the Cloud, and at the same time, many data-intensive scientific applications are being expressed as these services. In this paper, we explore an approach to accelerate service processing in a Cloud setting. We have developed a cooperative scheme for caching data output from services for reuse. We propose algorithms for scaling our cache system up during peak querying times, and back down to save costs. Using the Amazon EC2public Cloud, a detailed evaluation of our system has been performed, considering speed up and elastic scalability in terms resource allocation and relaxation.
David Chiu 0001, Apeksha Shetty, Gagan Agrawal
SC1
2009 Hierarchical Caches for Grid Workflows
abstract
From personal software to advanced systems, caching mechanisms have steadfastly been a ubiquitous means for reducing workloads. It is no surprise, then, that under the grid and cluster paradigms, middlewares and other large-scale applications often seek caching solutions. Among these distributed applications, scientific workflow management systems have gained ground towards mitigating the often painstaking process of composing sequences of scientific data sets and services to derive virtual data. In the past, workflow managers have relied on low-level system cache for reuse support. But in distributed query intensive environments, where high volumes of intermediate virtual data can potentially be stored anywhere on the grid, a novel cache structure is needed to efficiently facilitate workflow planning. In this paper, we describe an approach to combat the challenges of maintaining large, fast virtual data caches for workflow composition. A hierarchical structure is proposed for indexing scientific data with spatiotemporal annotations across grid nodes. Our experimental results show that our hierarchical index is scalable and outperforms a centralized indexing scheme by an exponential factor in query intensive environments.
David Chiu 0001, Gagan Agrawal
CCGRID1
2009 A Dynamic Approach toward QoS-Aware Service Workflow Composition
abstract
Web service-based workflow management systems have garnered considerable attention for automating and scheduling dependent operations. Such systems often support user preferences, e.g., time of completion, but with the rebirth of distributed computing via the grid/cloud, new challenges are abound: multiple disparate data sources, networks, nodes, and the potential for moving very large datasets. In this paper, we present a framework for integrating QoS support in a service workflow composition system. The relationship between workflow execution time and accuracy is exploited through an automatic workflow composition scheme. The algorithm, equipped with a framework for defining cost models on service completion times and error propagation, composes service workflows which can adapt to user's QoS preferences.
David Chiu 0001, Sagar Deshpande, Gagan Agrawal, Rongxing Li
ICWS1
2009 Enabling Ad Hoc Queries over Low-Level Scientific Data Sets
David Chiu 0001, Gagan Agrawal
SSDBM1
2008 Composing geoinformatics workflows with user preferences
abstract
With the advent of the data grid came a novel distributed scientific computing paradigm known as service-oriented sci-ence. Among the plethora of systems included under this framework are scientific workflow management systems, which enable large-scale process scheduling and execution. To en-sure quality of service, these systems typically seek to min-imize workflow execution time as well as costs for slices of data grid access. The geospatial domain, among other sci-ences, involves yet another optimization factor, the accu-racy of results. The relationship between execution time and workflow accuracy can often be exploited to offer more flexibility in handling user preferences. We present a system which meets user constraints through a dynamic adjustment of the accuracy of workflow results.
David Chiu 0001, Sagar Deshpande, Gagan Agrawal, Rongxing Li
GIS1