VLDB 2026 Research / reviewers in the wild / expert
Carl Kesselman
dblp:06/4477
· DBLP profile ↗
79ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0003-0917-1562ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSecurity and privacy · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | It's the Data StupidabstractArtificial Intelligence and Machine Learning have emerged as a promising approach to scientific investigations, but there is a persistent shortage of high-quality, properly annotated datasets suitable for training models. Here, we outline some of the widely reported characteristics for making AI-ready data and compare that with FAIR data. We discuss the limitations of traditional data repositories and the challenges associated with establishing a data repository that can grow with scientific communities and accommodate rapid evolution in research priorities. Finally, we introduce the SCALE principles for repository design that offer a proven framework for creating sustainable, scalable data repositories that can adapt to new data models and methodologies, ensuring that software infrastructure serves research needs rather than constraining them. Carl Kesselman, Robert Schuler |
eScience | 1 |
| 2025 | From Data to Decision: Data-Centric Infrastructure for Reproducible ML in Collaborative eScienceabstractReproducibility remains a central challenge in machine learning (ML), especially in collaborative eScience projects where teams iterate over data, features, and models. Current ML workflows are often dynamic yet fragmented, relying on informal data sharing, ad hoc scripts, and loosely connected tools. This fragmentation impedes transparency, reproducibility, and the adaptability of experiments over time. This paper introduces a data-centric framework for lifecycle-aware reproducibility, centered around six structured artifacts: Dataset, Feature, Workflow, Execution, Asset, and Controlled Vocabulary. These artifacts formalize the relationships between data, code, and decisions, enabling ML experiments to be versioned, interpretable, and traceable over time. The approach is demonstrated through a clinical ML use case of glaucoma detection, illustrating how the system supports iterative exploration, improves reproducibility, and preserves the provenance of collaborative decisions across the ML lifecycle. Carl Kesselman, Tran Huy Nguyen, Benjamin Yixing Xu, Kyle Bolo, Kimberley Yu |
eScience | 2 |
| 2024 | Deriva-ML: A Continuous FAIRness Approach to Reproducible Machine Learning ModelsabstractIncreasingly, artificial intelligence (AI) and machine learning (ML) are used in eScience applications [9]. While these approaches have great potential, the literature has shown that ML-based approaches frequently suffer from results that are either incorrect or unreproducible due to mismanagement or misuse of data used for training and validating the models [12], [15]. Recognition of the necessity of high-quality data for correct ML results has led to data-centric ML approaches that shift the central focus from model development to creation of high-quality data sets to train and validate the models [14], [20]. However, there are limited tools and methods available for data-centric approaches to explore and evaluate ML solutions for eScience problems which often require collaborative multidisciplinary teams working with models and data that will rapidly evolve as an investigation unfolds [1]. In this paper, we show how data management tools based on the principle that all of the data for ML should be findable, accessible, interoperable and reusable (i.e. FAIR [26]) can significantly improve the quality of data that is used for ML applications. When combined with best practices that apply these tools to the entire life cycle of an ML-based eScience investigation, we can significantly improve the ability of an eScience team to create correct and reproducible ML solutions. We propose an architecture and implementation of such tools and demonstrate through two use cases how they can be used to improve ML-based eScience investigations. Carl Kesselman, Mike D'Arcy, Michael Pazzani, Benjamin Yizing Xu |
e-Science | 2 |
| 2024 | Creating Thriving Data-Centric Communities from Basic Research to Commercial ApplicationsabstractThe ability to accumulate and analyze large quantities of data is rapidly becoming a competitive advantage not only in science but in the broader economy as well. Advances such as AlphaFold, the AI-based protein prediction tool and ChatGPT the large language model-based chat bot, have ignited enormous excitement in science and industry for leveraging data and computational techniques to solve important problems. However, what is typically lost in all the excitement is the fact that such startling achievements were only possible after a critical mass of high quality data existed to train models using machine learning algorithms. Both examples relied on open data sources that were generated painstakingly by user communities over the course of decades. We argue that in order to unlock future high impact data science achievements like these will require a culture of and skill set for data management, sharing and reuse. In this paper, we describe our work within the dental, oral, and craniofacial community to create such a data sharing community that has grown out of basic research to increasingly touch on clinical sciences and even commercial applications. Robert Schuler, Carl Kesselman |
e-Science | 2 |
| 2023 | Let's Put the Science in eScienceabstractThe underlying premise behind eScience is that computational methods and data-driven approaches can contribute to scientific discovery on a par with, or even superior to, traditional experimental methods; that the combination of computers, software, and extant data collections are the modern equivalent to the scientific instruments that have led to our understanding of fundamental laws in physics, chemistry, biology, and other domains. However, a robust methodology for making the results of eScience activities “scientific” is lacking, with significant consequences. In this brief paper we propose a shift in perspective as to what it means to create an eScience-based result and how the scientific validity of eScience experiments might be improved. Carl Kesselman, Robert Schuler, Ian T. Foster |
e-Science | 1 |
| 2023 | Database Evolution, by Scientists, for Scientists: A Case StudyabstractDatabase management systems have been used to great advantage for industry usage scenarios. As science becomes increasingly dependent on carefully organized and curated data to inform and drive new discoveries, the need for database management systems has grown significantly. The long standing “20 questions” method was advocated in early studies of applying relational databases for science in order to elicit requirements for designing and developing information systems for scientific data. It has been observed, however, that database designs become outdated within months of usage leading to degradation in the quality of the database schema. In addition, there is limited evidence that scientists themselves have the tools and processes necessary to develop and maintain scientific databases without reliance on database administrators. Beyond learning to query databases, scientists need tools to create and evolve databases and guidance on how to apply those tools to develop information systems. In this paper, we present a simplified methodology for database evolution for scientists and a case study of database evolution by a scientist in the context of a research database for cell modeling. We include a detailed analysis of the activities and processes employed by the scientist during the schema evolution. Our results show that a scientist can successfully evolve a complex information system driven by new research requirements. Robert Schuler, Jitin Singla, Brinda Vallat, Kate L. White, Helen M. Berman, Carl Kesselman |
e-Science | 6 |
| 2022 | Managing Database-Application Co-Evolution in a Scientific Data EcosystemabstractScientific databases used for organizing, archiving, collaborating and sharing research data depend on a well-defined schema to accurately reflect the scientific domain and on database-driven applications for supporting key user interactions with the database. Applications that interact with a database typically depend on some form of schema mappings, such as object-relational mappings, to inform the application of how to query and manipulate the database. The presence of schema mappings, however, further exacerbates the already difficult task of evolving the database schema. Database migration utilities provide some help by coordinating schema evolution scripts with application code changes, but only automate the simplest schema mapping changes. In this paper, we present an approach to coupled database-application evolution by extending a database evolution language with model management operations. We introduce a novel set of model management operations and define their semantics and then describe how they may be integrated into schema modification operators. We then present an evaluation of the concepts from real-world usage of model mappings in scientific database deployments. Robert Schuler, Carl Kesselman |
e-Science | 2 |
| 2021 | CHiSEL: a user-oriented framework for simplifing database evolution
Robert Schuler, Carl Kesselman |
Distributed Parallel Databases | 2 |
| 2020 | Towards Co-Evolution of Data-Centric EcosystemsabstractDatabase evolution is a notoriously difficult task, and it is exacerbated by the necessity to evolve database-dependent applications. As science becomes increasingly dependent on sophisticated data management, the need to evolve an array of database-driven systems will only intensify. In this paper, we present an architecture for data-centric ecosystems that allows the components to seamlessly co-evolve by centralizing the models and mappings at the data service and pushing model-adaptive interactions to the database clients. Boundary objects fill the gap where applications are unable to adapt and need a stable interface to interact with the components of the ecosystem. Finally, evolution of the ecosystem is enabled via integrated schema modification and model management operations. We present use cases from actual experiences that demonstrate the utility of our approach. Robert Schuler, Karl Czajkowski, Mike D'Arcy, Hongsuda Tangmunarunkit, Carl Kesselman |
SSDBM | 5 |
| 2019 | Toward FAIR Knowledge Turns in BioinformaticsabstractSharing of bioinformatics data within research communities holds the promise of facilitating more rapid discovery, yet the volume of data is growing at a pace exponentially greater than what traditional biocuration can support. We present here an approach that we have used to empower data producing researchers to curate high quality shared data that is ready for reuse and re-analysis. Robert Schuler, Alejandro Bugacov, Matthew Blow, Carl Kesselman |
BIBM | 4 |
| 2019 | A High-level User-oriented Framework for Database EvolutionabstractDatabases are well suited to the task of describing and organizing research datasets, however, the difficulties of using database management systems effectively have resulted in their limited usage among domain scientists. Scientists operate in an environment that is changing steadily with new experimental protocols, instruments, and discoveries that impact what datasets they generate and how they describe and organize them. In order to manage datasets for a scientific application, scientists need to routinely revise their database schemas to reflect these changes. Unfortunately, evolving a database is one of the well-known and most difficult aspects of database usage. The conventional data definition and manipulation languages offer relatively low-level programming abstractions to perform complex database evolution tasks, and therefore require specialized technical skills not possessed by most domain scientists. A simplified means of expressing database evolution operations can reduce the effort of keeping the scientific database in sync with changing requirements. This paper presents a high-level, user-oriented, schema evolution framework with an algebra of specialized schema modification operators. The approach allows introduction of novel operators as motivated by new requirements and is amenable to well established optimization techniques for efficient planning and execution. We present the framework and its implementation, and we demonstrate its utility in an exemplar use case and performance evaluation. Robert Schuler, Carl Kesselman |
SSDBM | 2 |
| 2018 | ERMrest: a web service for collaborative data managementabstractThe foundation of data oriented scientific collaboration is the ability for participants to find, access and reuse data created during the course of an investigation, what has been referred to as the FAIR principles. In this paper, we describe ERMrest, a collaborative data management service that promotes data oriented collaboration by enabling FAIR data management throughout the data life cycle. ERMrest is a RESTful web service that promotes discovery and reuse by organizing diverse data assets into a dynamic entity relationship model. We present details on the design and implementation of ERMrest, data on its performance and its use by a range of collaborations to accelerate and enhance their scientific output. Karl Czajkowski, Carl Kesselman, Robert Schuler, Hongsuda Tangmunarunkit |
SSDBM | 2 |
| 2018 | Towards an efficient and effective framework for the evolution of scientific databasesabstractDatabase systems are well suited to scientific data management and analysis workloads, however, a database must evolve to keep pace with changing requirements and adjust to changes in the domain conceptualization as applications mature. Evolving a database (i.e., updating its schema and instance data) is one of the greatest challenges in database maintenance and the difficulties are compounded by the lack of sufficient tools to support scientists. This paper presents a schema evolution framework based on an algebraic approach that introduces extended and higher-level composite relational operators tailored to the task of schema evolution. These higher-level operators simplify the task of evolving a database for non-expert users, while enabling efficient evaluation of schema evolution expressions. Robert Schuler, Carl Kesselman |
SSDBM | 2 |
| 2017 | Experiences with DERIVA: An Asset Management Platform for Accelerating eScienceabstractThe pace of discovery in eScience is increasingly dependent on a scientist's ability to acquire, curate, integrate, analyze, and share large and diverse collections of data. It is all too common for investigators to spend inordinate amounts of time developing ad hoc procedures to manage their data. In previous work, we presented Deriva, a Scientific Asset Management System, designed to accelerate data driven discovery. In this paper, we report on the use of Deriva in a number of substantial and diverse eScience applications. We describe the lessons we have learned, both from the perspective of the Deriva technology, as well as the ability and willingness of scientists to incorporate Scientific Asset Management into their daily workflows. Alejandro Bugacov, Karl Czajkowski, Carl Kesselman, Robert Schuler, Hongsuda Tangmunarunkit |
eScience | 3 |
| 2017 | ERMRest: A Collaborative Data Catalog with Fine Grain Access ControlabstractCreating and maintaining an accurate description of data assets and the relationships between assets is a critical aspect of making data findable, accessible, interoperable, and reusable (FAIR). Typically, such metadata are created and maintained in a data catalog by a curator as part of data publication. However, allowing metadata to be created and maintained by data producers as the data is generated rather then waiting for publication can have significant advantages in terms of productivity and repeatability. The responsibilities for metadata management need not fall on any one individual, but rather may be delegated to appropriate members of a collaboration, enabling participants to edit or maintain specific attributes, to describe relationships between data elements, or to correct errors. To support such collaborative data editing, we have created ERMrest, a relational data service for the Web that enables the creation, evolution and navigation of complex models used to describe and structure diverse file or relational data objects. A key capability of ERMrest is its ability to control operations down to the level of individual data elements, i.e. fine-grained access control, so that many different modes of data-oriented collaboration can be supported. In this paper we introduce ERMrest and describe its fine-grained access control capabilities that support collaborative editing. ERMrest is in daily use in many data driven collaborations and we describe a sample policy that is based on a common biocuration pattern. Karl Czajkowski, Carl Kesselman, Robert Schuler |
eScience | 2 |
| 2016 | Big Data Technologies for Biomedical Knowledge Discovery
Naveen Ashish, Arthur W. Toga, Ian T. Foster, Ivo D. Dinov, Carl Kesselman |
AMIA | 5 |
| 2016 | I'll take that to go: Big data bags and minimal identifiers for exchange of large, complex datasetsabstractBig data workflows often require the assembly and exchange of complex, multi-element datasets. For example, in biomedical applications, the input to an analytic pipeline can be a dataset consisting thousands of images and genome sequences assembled from diverse repositories, requiring a description of the contents of the dataset in a concise and unambiguous form. Typical approaches to creating datasets for big data workflows assume that all data reside in a single location, requiring costly data marshaling and permitting errors of omission and commission because dataset members are not explicitly specified. We address these issues by proposing simple methods and tools for assembling, sharing, and analyzing large and complex datasets that scientists can easily integrate into their daily workflows. These tools combine a simple and robust method for describing data collections (BDBags), data descriptions (Research Objects), and simple persistent identifiers (Minids) to create a powerful ecosystem of tools and services for big data analysis and sharing. We present these tools and use biomedical case studies to illustrate their use for the rapid assembly, sharing, and analysis of large datasets. Kyle Chard, Mike D'Arcy, Benjamin D. Heavner, Ian T. Foster, Carl Kesselman, Ravi K. Madduri, Alexis A. Rodriguez, Stian Soiland-Reyes, Carole A. Goble, Kristi Clark, Eric W. Deutsch, Ivo D. Dinov, Nathan D. Price 0001, Arthur W. Toga |
IEEE BigData | 5 |
| 2016 | Accelerating data-driven discovery with scientific asset managementabstractThe overhead and burden of managing data in complex discovery processes involving experimental protocols with numerous data-producing and computational steps has become the gating factor that determines the pace of discovery. The lack of comprehensive systems to capture, manage, organize and retrieve data throughout the discovery life cycle leads to significant overheads on scientists' time and effort, reduced productivity, lack of reproducibility, and an absence of data sharing. In “creative fields” like digital photography and music, digital asset management (DAM) systems for capturing, managing, curating and consuming digital assets like photos and audio recordings, have fundamentally transformed how these data are used. While asset management has not taken hold in eScience applications, we believe that transformation similar to that observed in the creative space could be achieved in scientific domains if appropriate ecosystems of asset management tools existed to capture, manage, and curate data throughout the scientific discovery process. In this paper, we introduce DERIVA, a framework and infrastructure for asset management in eScience and present initial results from its usage in active research use cases. Robert Schuler, Carl Kesselman, Karl Czajkowski |
eScience | 2 |
| 2015 | A system to build distributed multivariate models and manage disparate data sharing policies: implementation in the scalable national network for effectiveness researchabstractBACKGROUND: Centralized and federated models for sharing data in research networks currently exist. To build multivariate data analysis for centralized networks, transfer of patient-level data to a central computation resource is necessary. The authors implemented distributed multivariate models for federated networks in which patient-level data is kept at each site and data exchange policies are managed in a study-centric manner. OBJECTIVE: The objective was to implement infrastructure that supports the functionality of some existing research networks (e.g., cohort discovery, workflow management, and estimation of multivariate analytic models on centralized data) while adding additional important new features, such as algorithms for distributed iterative multivariate models, a graphical interface for multivariate model specification, synchronous and asynchronous response to network queries, investigator-initiated studies, and study-based control of staff, protocols, and data sharing policies. MATERIALS AND METHODS: Based on the requirements gathered from statisticians, administrators, and investigators from multiple institutions, the authors developed infrastructure and tools to support multisite comparative effectiveness studies using web services for multivariate statistical estimation in the SCANNER federated network. RESULTS: The authors implemented massively parallel (map-reduce) computation methods and a new policy management system to enable each study initiated by network participants to define the ways in which data may be processed, managed, queried, and shared. The authors illustrated the use of these systems among institutions with highly different policies and operating under different state laws. DISCUSSION AND CONCLUSION: Federated research networks need not limit distributed query functionality to count queries, cohort discovery, or independently estimated analytic models. Multivariate analyses can be efficiently and securely conducted without patient-level data transport, allowing institutions with strict local data storage requirements to participate in sophisticated analyses based on federated research networks. Daniella Meeker, Xiaoqian Jiang, Michael E. Matheny, Claudiu Farcas, Mike D'Arcy, Laura Pearlman, Lavanya Nookala, Michele E. Day, Katherine K. Kim, Hyeon-Eui Kim, Aziz A. Boxwala, Robert El-Kareh, Grace Kuo, Frederic S. Resnic, Carl Kesselman, Lucila Ohno-Machado |
J. Am. Medical Informatics Assoc. | 15 |
| 2015 | Big biomedical data as the key resource for discovery scienceabstractModern biomedical data collection is generating exponentially more data in a multitude of formats. This flood of complex data poses significant opportunities to discover and understand the critical interplay among such diverse domains as genomics, proteomics, metabolomics, and phenomics, including imaging, biometrics, and clinical data. The Big Data for Discovery Science Center is taking an "-ome to home" approach to discover linkages between these disparate data sources by mining existing databases of proteomic and genomic data, brain images, and clinical assessments. In support of this work, the authors developed new technological capabilities that make it easy for researchers to manage, aggregate, manipulate, integrate, and model large amounts of distributed data. Guided by biological domain expertise, the Center's computational resources and software will reveal relationships and patterns, aiding researchers in identifying biomarkers for the most confounding conditions and diseases, such as Parkinson's and Alzheimer's. Arthur W. Toga, Ian T. Foster, Carl Kesselman, Ravi K. Madduri, Kyle Chard, Eric W. Deutsch, Nathan D. Price 0001, Gwênlyn Glusman, Benjamin D. Heavner, Ivo D. Dinov, Joseph Ames, John D. Van Horn, Roger Kramer, Leroy E. Hood |
J. Am. Medical Informatics Assoc. | 3 |
| 2014 | Digital asset management for heterogeneous biomedical data in an era of data-intensive scienceabstractBiomedical research depends upon increasingly high throughput instruments and sophisticated data analytics. In spite of the significant overhead of handling research data, there is little support for researchers to manage and organize data for purposes of exploration, analysis, and ultimately publication. Shared file systems with metadata coded into directory hierarchies and spreadsheets are the common practice. In this paper, we present a digital asset management approach and system for streamlining data operations and reducing data management overheads for biomedical researchers. It consists of data management tasks including storage, archival, annotation, search, cataloging, publication, and collaboration. We present a preliminary performance evaluation of a key component of the system, and we demonstrate the utility of this approach in a pilot deployment and user study. Robert Schuler, Carl Kesselman, Karl Czajkowski |
BIBM | 2 |
| 2012 | A resiliency model for high performance infrastructure based on logical encapsulationabstractAn emerging trend in distributed systems is the creation of dynamically provisioned heterogeneous high performance platforms that include the co-allocation of both virtualized computing and network attached storage volumes offering NAS and SAN level data services. These high performance computing environments support parallel applications performing traditional file system operations. As with any parallel platform the ability to continue computation in the face of component failures is an important characteristic. Achieving resiliency in heterogeneous environments presents unique challenges and opportunities not found in homogeneous aggregations of computing resources. We present a logical encapsulation model for heterogeneous high performance infrastructure, which enables a reactive resiliency approach for federations of virtual machines and externally hosted physical storage volumes. Asynchronous state capture and restoration models are presented for individual resources, which are composed into non-blocking resiliency models for logical encapsulations. We perform an evaluation that demonstrates our methodology has greater overall flexibility and significant performance improvements when compared to current resiliency approaches in virtualized distributed execution environments. James J. Moore, Carl Kesselman |
HPDC | 2 |
| 2011 | Applications of the Pipeline Environment for Visual Informatics and Genomics ComputationsabstractBACKGROUND: Contemporary informatics and genomics research require efficient, flexible and robust management of large heterogeneous data, advanced computational tools, powerful visualization, reliable hardware infrastructure, interoperability of computational resources, and detailed data and analysis-protocol provenance. The Pipeline is a client-server distributed computational environment that facilitates the visual graphical construction, execution, monitoring, validation and dissemination of advanced data analysis protocols. RESULTS: This paper reports on the applications of the LONI Pipeline environment to address two informatics challenges - graphical management of diverse genomics tools, and the interoperability of informatics software. Specifically, this manuscript presents the concrete details of deploying general informatics suites and individual software tools to new hardware infrastructures, the design, validation and execution of new visual analysis protocols via the Pipeline graphical interface, and integration of diverse informatics tools via the Pipeline eXtensible Markup Language syntax. We demonstrate each of these processes using several established informatics packages (e.g., miBLAST, EMBOSS, mrFAST, GWASS, MAQ, SAMtools, Bowtie) for basic local sequence alignment and search, molecular biology data analysis, and genome-wide association studies. These examples demonstrate the power of the Pipeline graphical workflow environment to enable integration of bioinformatics resources which provide a well-defined syntax for dynamic specification of the input/output parameters and the run-time execution controls. CONCLUSIONS: The LONI Pipeline environment http://pipeline.loni.ucla.edu provides a flexible graphical infrastructure for efficient biomedical computing and distributed informatics research. The interactive Pipeline resource manager enables the utilization and interoperability of diverse types of informatics resources. The Pipeline client-server model provides computational power to a broad spectrum of informatics investigators--experienced developers and novice users, user with or without access to advanced computational-resources (e.g., Grid, data), as well as basic and translational scientists. The open development, validation and dissemination of computational networks (pipeline workflows) facilitates the sharing of knowledge, tools, protocols and best practices, and enables the unbiased validation and replication of scientific findings by the entire community. Ivo D. Dinov, Federica Torri, Fabio Macciardi, Petros Petrosyan, Alen Zamanyan, Paul R. Eggert, Jonathan Pierce, Alex Genco, James A. Knowles, Andrew P. Clark, John D. Van Horn, Joseph Ames, Carl Kesselman, Arthur W. Toga |
BMC Bioinform. | 14 |
| 2011 | Enabling collaborative research using the Biomedical Informatics Research Network (BIRN)abstractOBJECTIVE: As biomedical technology becomes increasingly sophisticated, researchers can probe ever more subtle effects with the added requirement that the investigation of small effects often requires the acquisition of large amounts of data. In biomedicine, these data are often acquired at, and later shared between, multiple sites. There are both technological and sociological hurdles to be overcome for data to be passed between researchers and later made accessible to the larger scientific community. The goal of the Biomedical Informatics Research Network (BIRN) is to address the challenges inherent in biomedical data sharing. MATERIALS AND METHODS: BIRN tools are grouped into 'capabilities' and are available in the areas of data management, data security, information integration, and knowledge engineering. BIRN has a user-driven focus and employs a layered architectural approach that promotes reuse of infrastructure. BIRN tools are designed to be modular and therefore can work with pre-existing tools. BIRN users can choose the capabilities most useful for their application, while not having to ensure that their project conforms to a monolithic architecture. RESULTS: BIRN has implemented a new software-based data-sharing infrastructure that has been put to use in many different domains within biomedicine. BIRN is actively involved in outreach to the broader biomedical community to form working partnerships. CONCLUSION: BIRN's mission is to provide capabilities and services related to data sharing to the biomedical research community. It does this by forming partnerships and solving specific, user-driven problems whose solutions are then available for use by other groups. Karl G. Helmer, José Luis Ambite, Joseph Ames, Rachana Ananthakrishnan, Gully A. P. C. Burns, Ann L. Chervenak, Ian T. Foster, Liming Lee, David B. Keator, Fabio Macciardi, Ravi K. Madduri, John-Paul Navarro, Steven G. Potkin, Bruce R. Rosen, Seth Ruffins, Robert Schuler, Jessica A. Turner, Arthur W. Toga, Christina Williams, Carl Kesselman |
J. Am. Medical Informatics Assoc. | 20 |
| 2009 | The Globus Replica Location Service: Design and ExperienceabstractDistributed computing systems employ replication to improve overall system robustness, scalability, and performance. A replica location service (RLS) offers a mechanism to maintain and provide information about physical locations of replicas. This paper defines a design framework for RLSs that supports a variety of deployment options. We describe the RLS implementation that is distributed with the Globus toolkit and is in production use in several grid deployments. Features of our modular implementation include the use of soft-state protocols to populate a distributed index and Bloom filter compression to reduce overheads for distribution of index information. Our performance evaluation demonstrates that the RLS implementation scales well for individual servers with millions of entries and up to 100 clients. We describe the characteristics of existing RLS deployments and discuss how RLS has been integrated with higher-level data management services. Ann L. Chervenak, Robert Schuler, Matei Ripeanu, Muhammad Ali Amer, Shishir Bharathi, Ian T. Foster, Adriana Iamnitchi, Carl Kesselman |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2008 | Grid Resource Abstraction, Virtualization, and Provisioning for Time-Targeted ApplicationsabstractAs a variety of science applications are integrated with large-scale HPDC (high performance distributed computing) technologies, timely resource allocation is revealed as a critical requirement to be considered. This paper introduces a new HPDC resource management paradigm named resource slot which defines a network of logical machines across time and space. A resource slot is not only a resource programming target but also a virtualized resource provisioning framework for a variety of resource management paradigms by encapsulating the resource management complexity. Especially, we present a resource provisioning technique named guided redundant submission (GRS), which probabilistically guarantees a timely resource slot allocation. Experimental results performed against 8 clusters in production show that about 5 redundant resources per slot can secure slot allocation with up to 36 logical machines, each cluster having an availability probability as low as 0.25 and the target success probability of slot allocation is 0.95. Yang-Suk Kee, Carl Kesselman |
CCGRID | 2 |
| 2008 | Enabling personal clusters on demand for batch resources using commodity softwareabstractProviding QoS (quality of service) in batch resources against the uncertainty of resource availability due to the space-sharing nature of scheduling policies is a critical capability required for high-performance computing. This paper introduces a technique called personal cluster which reserves a partition of batch resources on user's demand in a best-effort manner. A personal cluster provides a private cluster dedicated to the user during a user-specified time period by installing a user-level resource manager on the resource partition. This technique not only enables cost-effective resource utilization and efficient task management but also provides the user a uniform interface to heterogeneous resources regardless of local resource management software. A prototype implementation using a PBS batch resource manager and Globus Toolkits based on Web services shows that the overhead of instantiating a personal cluster of medium size is small, which is just about 1 minute for a personal cluster having 32 processors. Yang-Suk Kee, Carl Kesselman, Daniel Nurmi, Richard Wolski |
IPDPS | 2 |
| 2008 | Virtual Organizations By the RulesabstractIncreasingly, collaborative activities in science are using the concept of virtual organization as an organizing principle. One benefit of viewing these collaborations from an organizational perspective is that there is a long history of studying how organizations can be structured to function effectively. Many of these organizational principles have been reflected in the design of enterprise architectures and the use of service oriented architecture concepts as an implementation vehicle for capturing these organizational constructs. One approach to meeting organizational requirements in systems architecture has been to express organizational structure in terms of business roles, business processes and business rules. To date however, this type of analysis and associated infrastructure tools has not been applied in any consistent way to the concept of virtual organizations and their associated scientific applications. In this talk, the author explores these established approaches to business IT systems and their applicability to the virtual organizations that are being created to support scientific endeavors. As an example, I will describes how data management policies for virtual organization can be expressed as business rules, and implemented via existing business rules engines. Carl Kesselman |
PDCAT | 1 |
| 2007 | A provisioning model and its comparison with best-effort for performance-cost optimization in gridsabstractThe resource availability in Grids is generally unpredictable due to the autonomous and shared nature of the Grid resources and stochastic nature of the workload resulting in a best effort quality of service. The resource providers optimize for throughput and utilization whereas the users optimize for application performance. We present a cost-based model where the providers advertise resource availability to the user community. We also present a multi-objective genetic algorithm formulation for selecting the set of resources to be provisioned that optimizes the application performance while minimizing the resource costs. We use trace-based simulations to compare the application performance and cost using the provisioned and the best effort approach with a number of artificially generated workflow-structured applications and a seismic hazard application from the earthquake science community. The provisioned approach shows promising results when the resources are under high utilization and/or the applications have significant resource requirements. Carl Kesselman, Ewa Deelman |
HPDC | 2 |
| 2006 | Managing Large-Scale Workflow Execution from Resource Provisioning to Provenance Tracking: The CyberShake ExampleabstractThis paper discusses the process of building an environment where large-scale, complex, scientific analysis can be scheduled onto a heterogeneous collection of computational and storage resources. The example application is the Southern California Earthquake Center (SCEC) CyberShake project, an analysis designed to compute probabilistic seismic hazard curves for sites in the Los Angeles area. We explain which software tools were used to build to the system, describe their functionality and interactions. We show the results of running the CyberShake analysis that included over 250,000 jobs using resources available through SCEC and the TeraGrid. Ewa Deelman, Scott Callaghan, Edward Field, Hunter Francoeur, Robert Graves, Vipin Gupta, Thomas H. Jordan, Carl Kesselman, Philip Maechling, John Mehringer, Gaurang Mehta, David Okaya, Karan Vahi |
e-Science | 9 |
| 2006 | Dynamic Infrastructure for Systems Level Science
Carl Kesselman |
e-Science | 1 |
| 2006 | Application-Level Resource Provisioning on the GridabstractIn this paper, we present algorithms for Grid resource provisioning that employ agreement-based resource management. These algorithms allow userlevel resource allocation and scheduling of applications that are structured as a precedenceconstrained set of tasks. We present a provisioning model where the resource availability in the Grid can be enumerated as a set of slots. A slot is defined as a number of processors available from a certain start time for a certain duration at a certain cost. Using a cost model that combines the cost of resource allocation and the expected application runtime, we evaluate the performance of the Min-Min and of the Genetic algorithm (GA)-based heuristics for a range of synthetic applications. We show that the GA paired with a list scheduling algorithm can obtain significantly better solutions than the Min-Min heuristic alone. Carl Kesselman, Ewa Deelman |
e-Science | 2 |
| 2006 | What makes workflows work in an opportunistic environment?abstractAbstract In this paper, we examine the issues of workflow mapping and execution in opportunistic environments such as the Grid. As applications become ever more complex, the process of choosing the appropriate resources and successfully executing the application components becomes ever more difficult. This may include extension or reduction of the initial workflow mapping as necessary for the actual execution. In this paper, we focus on the interplay between a workflow‐mapping component that plans the high‐level resource assignments and the workflow executor that oversees the component execution. We concentrate particularly on issues of data management and we draw from the experiences with mapping and execution systems: Pegasus, DAGMan and Stork. Copyright © 2005 John Wiley & Sons, Ltd. Ewa Deelman, Tevfik Kosar, Carl Kesselman, Miron Livny |
Concurr. Comput. Pract. Exp. | 3 |
| 2005 | Optimizing Grid-Based Workflow Execution
Carl Kesselman, Ewa Deelman |
J. Grid Comput. | 2 |
| 2005 | The Earth System Grid: Supporting the Next Generation of Climate Modeling ResearchabstractUnderstanding the Earth's climate system and how it might be changing is a preeminent scientific challenge. Global climate models are used to simulate past, present, and future climates, and experiments are executed continuously on an array of distributed supercomputers. The resulting data archive, spread over several sites, currently contains upwards of 100 TB of simulation data and is growing rapidly. Looking toward mid-decade and beyond, we must anticipate and prepare for distributed climate research data holdings of many petabytes. The Earth System Grid (ESG) is a collaborative interdisciplinary project aimed at addressing the challenge of enabling management, discovery, access, and analysis of these critically important datasets in a distributed and heterogeneous computational environment. The problem is fundamentally a Grid problem. Building upon the Globus toolkit and a variety of other technologies, ESG is developing an environment that addresses authentication, authorization for data access, large-scale data transport and management, services and abstractions for high-performance remote data access, mechanisms for scalable data replication, cataloging with rich semantic and syntactic information, data discovery, distributed monitoring, and Web-based portals for using the system. David E. Bernholdt, Shishir Bharathi, David Brown 0006, Kasidit Chanchio, Meili Chen, Ann L. Chervenak, Luca Cinquini, Bob Drach, Ian T. Foster, Peter Fox 0001, José García 0004, Carl Kesselman, Rob S. Markel, Don Middleton, Veronika Nefedova, Line C. Pouchard, Arie Shoshani, Alex Sim, Gary Strand, Dean N. Williams |
Proc. IEEE | 12 |
| 2005 | Agreement-Based Resource ManagementabstractOne of the criteria for the Grid infrastructure is the ability to share resources with nontrivial qualities of service. However, sharing resources in Grids is complicated in that is requires the ability bridge the differing policy requirements of the resource owners to create a consistent cross-organizational policy domain that delivers the necessary capability to the end user while respecting the policy requirements of the resource owner. Further complicating the management of Grid resources is the need to coordinate resource usage, the diversity of resource types and the variety of different management modes that may be used. We present a unifying resource management framework in which we can address these issues. The fundamental underlying concept in this framework is the representation of various resource management activities in terms of an agreement. Agreements abstract local management policy by representing an underlying resource strictly in terms of policy terms which it is willing to assert, and in doing so provides the basis for building a variety of alternative Grid resource management strategies. We introduce the concepts of agreement based resource management. We present a general agreement model and examine current resource management systems in the context of this model. We then discuss how agreement based resource management is being used as the basis for standards activities and next generation resource management services. Karl Czajkowski, Ian T. Foster, Carl Kesselman |
Proc. IEEE | 3 |
| 2004 | Performance and Scalability of a Replica Location Service
Ann L. Chervenak, Naveen Palavalli, Shishir Bharathi, Carl Kesselman, Robert Schwartzkopf |
HPDC | 4 |
| 2004 | The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon |
HPDC | 86 |
| 2004 | Distributed Hybrid Earthquake Engineering Experiments: Experiences with a Ground-Shaking Grid Application
Laura Pearlman, Carl Kesselman, Sridhar Gullapalli, B. F. Spencer Jr., Joe Futrelle, Kathleen Ricker, Ian T. Foster, Paul Hubbard, Charles R. Severance |
HPDC | 2 |
| 2004 | Grid-Based Metadata Services
Ewa Deelman, Malcolm P. Atkinson 0001, Ann L. Chervenak, Neil P. Chue Hong, Carl Kesselman, Sonal Patil, Laura Pearlman, Mei-Hui Su |
SSDBM | 6 |
| 2004 | Applications of Intelligent Agent Technology to The Grid
Carl Kesselman |
Web Intelligence | 1 |
| 2003 | An Ontology for Scientific Information in a Grid Environment: the Earth System GridabstractIn the emerging world of Grid Computing, shared computational, data, other distributed resources are becoming available to enable scientific advancement through collaborative research and collaboratories. This paper describes the increasing role of ontologies in the context of Grid Computing for obtaining, comparing and analyzing data. We present ontology entities and a declarative model that provide the outline for an ontology of scientific information. Relationships between concepts are also given. The implementation of some concepts described in this ontology is discussed within the context of the Earth System Grid II (ESG)[1]. Line C. Pouchard, Luca Cinquini, Bob Drach, Don Middleton, David E. Bernholdt, Kasidit Chanchio, Ian T. Foster, Veronika Nefedova, David Brown 0006, Peter Fox 0001, José García 0004, Gary Strand, Dean N. Williams, Ann L. Chervenak, Carl Kesselman, Arie Shoshani, Alex Sim |
CCGRID | 15 |
| 2003 | GridWorkflow: A Flexible Failure Handling Framework for the GridabstractThe generic, heterogeneous, and dynamic nature of the grid requires a new from of failure recovery mechanism to address its unique requirements such as support for diverse failure handling strategies, separation of failure handling strategies from application codes, and user-defined exception handling. We here propose a grid workflow system (grid-WFS), a flexible failure handling framework for the grid, which addresses these grid-unique failure recovery requirements. Central to the framework is flexibility by the use of workflow structure as a high-level recovery policy specification. We show how this use of high-level workflow structure allows users to achieve failure recovery in a variety of ways depending on the requirements and constraints of their applications. We also demonstrate that this use of workflow structure enables users to not only rapidly prototype and investigate failure handling strategies, but also easily change them by simply modifying the encompassing workflow structure, while the application code remains intact. Finally, we present an experimental evaluation of our framework using a simulation, demonstrating the value of supporting multiple failure recovery techniques in grid systems to achieve high performance in the presence of failures. Soonwook Hwang, Carl Kesselman |
HPDC | 2 |
| 2003 | Security for Grid ServicesabstractGrid computing is concerned with the sharing and coordinated use of diverse resources in distributed "virtual organizations." The dynamic and multiinstitutional nature of these environments introduces challenging security issues that demand new technical approaches. In particular, one must deal with diverse local mechanisms, support dynamic creation of services, and enable dynamic creation of trust domains. We describe how these issues are addressed in two generations of the Globus Toolkit/spl reg/. First, we review the Globus Toolkit version 2 (GT2) approach; then we describe new approaches developed to support the Globus Toolkit version 3 (GT3) implementation of the Open Grid Services Architecture, an initiative that is recasting Grid concepts within a service-oriented framework based on Web services. GT3's security implementation uses Web services security mechanisms for credential exchange and other purposes, and introduces a tight least-privilege model that avoids the need for any privileged network service. Von Welch, Frank Siebenlist, Ian T. Foster, John Bresnahan, Karl Czajkowski, Jarek Gawor, Carl Kesselman, Sam Meder, Laura Pearlman, Steven Tuecke |
HPDC | 7 |
| 2003 | Transparent Grid Computing: A Knowledge-Based Approach
Jim Blythe, Ewa Deelman, Yolanda Gil, Carl Kesselman |
IAAI | 4 |
| 2003 | Grid-Based Galaxy Morphology Analysis for the National Virtual ObservatoryabstractAs part of the development of the National Virtual Observatory (NVO), a Data Grid for astronomy, we have developed a prototype science application to explore the dynamical history of galaxy clusters by analyzing the galaxies' morphologies. The purpose of the prototype is to investigate how Grid-based technologies can be used to provide specialized computational services within the NVO environment. In this paper we focus on the key enabling technology components, particularly Chimera and Pegasus which are used to create and manage the computational workflow that must be present to deal with the challenging application requirements. We illustrate how the components interplay with each other and can be driven from a special purpose application portal. Ewa Deelman, Raymond Plante, Carl Kesselman, Mei-Hui Su, Gretchen Greene, Robert J. Hanisch, Niall Gaffney, Antonio Volpicelli, James Annis, Vijay Sekhri, Tamás Budavári, María A. Nieto-Santisteban, William O'Mullane, David Bohlender, Tom McGlynn, Arnold H. Rots, Olga Pevunova |
SC | 3 |
| 2003 | A Metadata Catalog Service for Data Intensive ApplicationsabstractAdvances in computational, storage and network technologies as well as middle ware such as the Globus Toolkit allow scientists to expand the sophistication and scope of data-intensive applications. These applications produce and analyze terabytes and petabytes of data that are distributed in millions of files or objects. To manage these large data sets efficiently, metadata or descriptive information about the data needs to be managed. There are various types of metadata, and it is likely that a range of metadata services will exist in Grid environments that are specialized for particular types of metadata cataloguing and discovery. In this paper, we present the design of a Metadata Catalog Service (MCS) that provides a mechanism for storing and accessing descriptive metadata and allows users to query for data items based on desired attributes. We describe our experience in using the MCS with several applications and present a scalability study of the service. Shishir Bharathi, Ann L. Chervenak, Ewa Deelman, Carl Kesselman, Mary Manohar, Sonal Patil, Laura Pearlman |
SC | 5 |
| 2003 | Ontology-Based Resource Matching in the Grid - The Grid Meets the Semantic Web
Hongsuda Tangmunarunkit, Stefan Decker, Carl Kesselman |
ISWC | 3 |
| 2003 | Multi-wavelength image space: another Grid-enabled scienceabstractAbstract We describe how the Grid enables new research possibilities in astronomy through multi‐wavelength images. To see sky images in the same pixel space, they must be projected to that space, a computer‐intensive process. There is thus a virtual data space induced that is defined by an image and the applied projection. This virtual data can be created and replicated with Planners and Replica catalog technology developed under the GriPhyN project. We plan to deploy our system (MONTAGE) on the U.S. Teragrid. Grid computing is also needed for ingesting data—computing background correction on each image—which forms a separate virtual data space. Multi‐wavelength images can be used for pushing source detection and statistics by an order of magnitude from current techniques; for optimization of multi‐wavelength image registration for detection and characterization of extended sources; and for detection of new classes of essentially multi‐wavelength astronomical phenomena. The paper discusses both the Grid architecture and the scientific goals. Copyright © 2003 John Wiley & Sons, Ltd. Roy Williams, G. Bruce Berriman, Ewa Deelman, John Good, Joseph C. Jacob, Carl Kesselman, Carol Lonsdale, Seb Oliver, Thomas A. Prince |
Concurr. Comput. Pract. Exp. | 6 |
| 2003 | Mapping Abstract Complex Workflows onto Grid Environments
Ewa Deelman, Jim Blythe, Yolanda Gil, Carl Kesselman, Gaurang Mehta, Karan Vahi, Kent Blackburn, Albert Lazzarini, Adam Arbree, Richard Cavanaugh, Scott Koranda |
J. Grid Comput. | 4 |
| 2003 | A Flexible Framework for Fault Tolerance in the Grid
Soonwook Hwang, Carl Kesselman |
J. Grid Comput. | 2 |
| 2003 | High-performance remote access to climate simulation data: a challenge problem for data grid technologies
Ann L. Chervenak, Ewa Deelman, Carl Kesselman, William E. Allcock, Ian T. Foster, Veronika Nefedova, Jason Lee 0001, Alex Sim, Arie Shoshani, Bob Drach, Dean N. Williams, Don Middleton |
Parallel Comput. | 3 |
| 2002 | Dependability and the Grid: Issues and Challenges
Richard D. Schlichting, Andrew A. Chien, Carl Kesselman, Keith Marzullo, James S. Plank, Santosh K. Shrivastava |
DSN | 3 |
| 2002 | GriPhyN and LIGO, Building a Virtual Data Grid for Gravitational Wave ScientistsabstractMany Physics experiments today generate large volumes of data. That data is then processed in a variety of ways in order to achieve the understanding of fundamental physical phenomena. The goal of the NSF-funded GriPhyN project (Grid Physics Network) is to enable scientists to seamlessly access data whether it is raw experimental data or a data product which is a result of further processing. GriPhyN provides a new degree of transparency in how data-handling and processing capabilities are integrated to deliver data products to end-users or applications, so that requests for such products are easily mapped into computation and/or data access at multiple locations. GriPhyN refers to the set of all data products available to the user as virtual data. Among the physics applications participating in the project is the Laser Interferometer Gravitational-wave Observatory (LIGO), which is being built to observe the gravitational waves predicted by general relativity. We describe our initial design and prototype of a virtual data Grid for LIGO. Ewa Deelman, Carl Kesselman, Gaurang Mehta, Leila Meshkat, Laura Pearlman, Kent Blackburn, Phil Ehrens, Albert Lazzarini, Roy Williams, Scott Koranda |
HPDC | 2 |
| 2002 | SNAP: A Protocol for Negotiating Service Level Agreements and Coordinating Resource Management in Distributed Systems
Karl Czajkowski, Ian T. Foster, Carl Kesselman, Volker Sander, Steven Tuecke |
JSSPP | 3 |
| 2002 | Giggle: a framework for constructing scalable replica location servicesabstractIn wide area computing systems, it is often desirable to create remote read-only copies (replicas) of files. Replication can be used to reduce access latency, improve data locality, and/or increase robustness, scalability and performance for distributed applications. We define a replica location service (RLS) as a system that maintains and provides access to information about the physical locations of copies. An RLS typically functions as one component of a data grid architecture. This paper makes the following contributions. First, we characterize RLS requirements. Next, we describe a parameterized architectural framework, which we name Giggle (for GIGa-scale Global Location Engine), within which a wide range of RLSs can be defined. We define several concrete instantiations of this framework with different performance characteristics. Finally, we present initial performance results for an RLS prototype, demonstrating that RLS systems can be constructed that meet performance goals. Ann L. Chervenak, Ewa Deelman, Ian T. Foster, Leanne Guy, Wolfgang Hoschek, Adriana Iamnitchi, Carl Kesselman, Peter Z. Kunszt, Matei Ripeanu, Robert Schwartzkopf, Heinz Stockinger, Kurt Stockinger, Brian Tierney |
SC | 7 |
| 2002 | The Grid, Grid Services and the Semantic Web: Technologies and Opportunities
Carl Kesselman |
ISWC | 1 |
| 2002 | Data management and transfer in high-performance computational grid environments
William E. Allcock, Joseph Bester, John Bresnahan, Ann L. Chervenak, Ian T. Foster, Carl Kesselman, Sam Meder, Veronika Nefedova, Darcy Quesnel, Steven Tuecke |
Parallel Comput. | 6 |
| 2001 | Practical Resource Management for Grid-Based Visual ExplorationabstractComputational grids are enabling collaboration between scientists and organizations to generate and archive extremely large datasets across shared, distributed resources. There is a need to visually explore such data throughout the life-cycle of projects. Practical exploration of large datasets requires visualization tools that can function in the same grid environment in which the data is created and stored. Resource management interfaces are an important structural component of grid computing environments because they enable uniform access to the wide variety of resources necessary for scientific work. We describe a new advance-reservation system for graphics resources; and an application of existing grid technology to create general-purpose active storage systems. We report our experience with prototype infrastructure and application components, involving experiments coupling end-to-end resources for interactive visual exploration of large data in representative distributed environments. Karl Czajkowski, Alper K. Demir, Carl Kesselman, Marcus Thiébaux |
HPDC | 3 |
| 2001 | Grid Information Services for Distributed Resource SharingabstractGrid technologies enable large-scale sharing of resources within formal or informal consortia of individuals and/or institutions: what are sometimes called virtual organizations. In these settings, the discovery, characterization, and monitoring of resources, services, and computations are challenging problems due to the considerable diversity; large numbers, dynamic behavior, and geographical distribution of the entities in which a user might be interested. Consequently, information services are a vital part of any Grid software infrastructure, providing fundamental mechanisms for discovery and monitoring, and hence for planning and adapting application behavior. We present an information services architecture that addresses performance, security, scalability, and robustness requirements. Our architecture defines simple low-level enquiry and registration protocols that make it easy to incorporate individual entities into various information structures, such as aggregate directories that support a variety of different query languages and discovery strategies. These protocols can also be combined with other Grid protocols to construct additional higher-level services and capabilities such as brokering, monitoring, fault detection, and troubleshooting. Our architecture has been implemented as MDS-2, which forms part of the Globus Grid toolkit and has been widely deployed and applied. Karl Czajkowski, Carl Kesselman, Steven Fitzgerald, Ian T. Foster |
HPDC | 2 |
| 2001 | High-performance remote access to climate simulation data: a challenge problem for data grid technologiesabstractIn numerous scientific disciplines, terabyte and soon petabyte-scale data collections are emerging as critical community resources. A new class of Data Grid infrastructure is required to support management, transport, distributed access to, and analysis of these datasets by potentially thousands of users. Researchers who face this challenge include the Climate Modeling community, which performs long-duration computations accompanied by frequent output of very large files that must be further analyzed. We describe the Earth System Grid prototype, which brings together advanced analysis, replica management, data transfer, request management, and other technologies to support high-performance, interactive analysis of replicated data. We present performance results that demonstrate our ability to manage the location and movement of large datasets from the user's desktop. We report on experiments conducted over SciNET at SC'2000, where we achieved peak performance of 1.55Gb/s and sustained performance of 512.9Mb/s for data transfers between Texas and California. William E. Allcock, Ian T. Foster, Veronika Nefedova, Ann L. Chervenak, Ewa Deelman, Carl Kesselman, Jason Lee 0001, Alex Sim, Arie Shoshani, Bob Drach, Dean N. Williams |
SC | 6 |
| 2001 | Generalized Communicators in the Message Passing InterfaceabstractWe propose extensions to the message passing interface (MPI) that generalize the MPI communicator concept to allow multiple communication endpoints per process, dynamic creation of endpoints, and the transfer of endpoints between processes. The generalized communicator construct can be used to express a wide range of interesting communication structures, including collective communication operations involving multiple threads per process, communications between dynamically created threads or processes, and object-oriented applications in which communications are directed to specific objects. Furthermore, this enriched functionality can be provided in a manner that preserves backward compatibility with MPI. We describe the proposed extensions, illustrate their use with examples, and describe a prototype implementation in the popular MPI implementation MPICH. Erik D. Demaine, Ian T. Foster, Carl Kesselman, Marc Snir |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2000 | The data grid: Towards an architecture for the distributed management and analysis of large scientific datasets
Ann L. Chervenak, Ian T. Foster, Carl Kesselman, Charles Salisbury, Steven Tuecke |
J. Netw. Comput. Appl. | 3 |
| 1999 | Resource Co-Allocation in Computational GridsabstractApplications designed to execute on "computational grids" frequently require the simultaneous co-allocation of multiple resources in order to meet performance requirements. For example, several computers and network elements may be required in order to achieve real-time reconstruction of experimental data, while a large numerical simulation may require simultaneous access to multiple supercomputers. Motivated by these concerns, we have developed a general resource management architecture for Grid environments, in which resource co-allocation is an integral component. We examine the co-allocation problem in detail and present mechanisms that allow an application to guide resource selection during the co-allocation process; these mechanisms address issues relating to the allocation, monitoring, control, and configuration of distributed computations. We describe the implementation of co-allocators based on these mechanisms and present the results of microbenchmark studies and large-scale application experiments that provide insights into the costs and practical utility of our techniques. Karl Czajkowski, Ian T. Foster, Carl Kesselman |
HPDC | 3 |
| 1999 | A Network Performance Tool for Grid EnvironmentsabstractIn grid computing environments, network bandwidth discovery and allocation is a serious issue.Before their applications are running, grid users will need to choose hosts based on available bandwidth.Running applications may need to adapt to a changing set of hosts.Hence, a tool is needed for monitoring network performance that is integral to the grid environment.To address this need, Gloperf was developed as part of the Globus grid computing toolkit.Gloperf is designed for ease of deployment and makes simple, end-to-end TCP measurements requiring no special host permissions.Scalability is addressed by a hierarchy of measurements based on group membership and by limiting overhead to a small, acceptable, fixed percentage of the available bandwidth.Since this fixed overhead may push host-pair revisit time into the tens-of-hours, we also quantitatively examine the "trajectory" of the cost-error trade-off for measurement frequency. Craig A. Lee, James Stepanek, Richard Wolski, Carl Kesselman, Ian T. Foster |
SC | 4 |
| 1999 | The Globus project: a status report
Ian T. Foster, Carl Kesselman |
Future Gener. Comput. Syst. | 2 |
| 1998 | A Security Architecture for Computational GridsabstractState-of-the-artand emerging scientific applications require fast access to large quantities of data and commensurately fast computational resources.Both resources and data are oflen distributed in a wide-area network with components administered locally and independently.Computations may involve hundreds of processes that must be able to acquire resources dynamically and communicate efficiently.This paper analyzes the unique security requirements of large-scale distributed (grid) computing and develops a security policy and a corresponding security architecture.An implementation of the architecture within the Globus metacomputing toolkit is discussed. Ian T. Foster, Carl Kesselman, Gene Tsudik, Steven Tuecke |
CCS | 2 |
| 1998 | Application Experiences with the Globus ToolkitabstractThe development of applications and tools for high-performance "computational grids" is complicated by the heterogeneity and frequently dynamic behavior of the underlying resources; by the complexity of the applications themselves, which often combine aspects of supercomputing and distributed computing; and by the need to achieve high levels of performance. The Globus toolkit has been developed with the goal of simplifying this application development task, by providing implementations of various core services deemed essential for high-performance distributed computing. In this paper, we describe two large applications developed with this toolkit: a distributed interactive simulation and a teleimmersion system. We describe the process used to develop the applications, review the lessons learned and draw conclusions regarding the effectiveness of the toolkit approach. Sharon Brunett, Karl Czajkowski, Steven Fitzgerald, Ian T. Foster, Andrew E. Johnson 0001, Carl Kesselman, Jason Leigh, Steven Tuecke |
HPDC | 6 |
| 1998 | A Fault Detection Service for Wide Area Distributed ComputationsabstractThe potential for faults in distributed computing systems is a significant complicating factor for application developers. While a variety of techniques exist for detecting and correcting faults, the implementation of these techniques in a particular context can be difficult. Hence, we propose a fault detection service designed to be incorporated, in a modular fashion, into distributed computing systems, tools, or applications. This service uses well-known techniques based on unreliable fault detectors to detect and report component failure, while allowing the user to tradeoff timeliness of reporting against false positive rates. We describe the architecture of this service, report on experimental results that quantify its cost and accuracy, and describe its use in two applications, monitoring the status of system components of the GUSTO computational grid testbed and as part of the NetSolve network-enabled numerical solver. Paul Stelling, Ian T. Foster, Carl Kesselman, Craig A. Lee, Gregor von Laszewski |
HPDC | 3 |
| 1998 | A Resource Management Architecture for Metacomputing Systems
Karl Czajkowski, Ian T. Foster, Nicholas T. Karonis, Carl Kesselman, Stuart Martin, Warren Smith, Steven Tuecke |
JSSPP | 4 |
| 1997 | A Directory Service for Configuring High-Performance Distributed ComputationsabstractHigh-performance execution in distributed computing environments often requires careful selection and configuration not only of computers, networks, and other resources but also of the protocols and algorithms used by applications. Selection and configuration in turn require access to accurate, up-to-date information on the structure and state of available resources. Unfortunately, no standard mechanism exists for organizing or accessing such information. Consequently, different tools and applications adopt ad hoc mechanisms, or they compromise their portability and performance by using default configurations. We propose a solution to this problem: a Metacomputing Directory Service that provides efficient and scalable access to diverse, dynamic, and distributed information about resource structure and state. We define an extensible data model to represent the information required for distributed computing, and we present a scalable, high-performance, distributed implementation. The dat... Steven Fitzgerald, Ian T. Foster, Carl Kesselman, Gregor von Laszewski, Warren Smith, Steven Tuecke |
HPDC | 3 |
| 1997 | A Secure Communications Infrastructure for High-Performance Distributed ComputingabstractApplications that use high-speed networks to connect geographically distributed supercomputers, databases, and scientific instruments may operate over open networks and access valuable resources. Hence, they can require mechanisms for ensuring integrity and confidentiality of communications and for authenticating both users and resources. Security solutions developed for traditional client-server applications do not provide direct support for the program structures, programming tools, and performance requirements encountered in these applications. We address these requirements via a security-enhanced version of the Nexus communication library, which we use to provide secure versions of parallel libraries and languages, including the Message Passing Interface. These tools permit a fine degree of control over what, where, and when security mechanisms are applied. In particular, a single application can mix secure and nonsecure communication allowing the programmer to make fine-grained security/performance tradeoffs. We present performance results that quantify the performance of our infrastructure. Ian T. Foster, Nicholas T. Karonis, Carl Kesselman, Gregory A. Koenig, Steven Tuecke |
HPDC | 3 |
| 1997 | Evaluating the Performance Limitations of MPMD CommunicationabstractThe MPMD approach for parallel computing is attractive for programmers who seek fast development cycles, high code re-use, and modular programming, or whose applications exhibit irregular computation loads and communication patterns. RPC is widely adopted as the communication abstraction for crossing address space boundaries. However, the communication overheads of existing RPC-based systems are usually an order of magnitude higher than those found in highly tuned SPMD systems. This problem has thus far limited the appeal of high-level programming languages based on MPMD models in the parallel computing community.This paper investigates the fundamental limitations of MPMD communication using a case study of two parallel programming languages, Compositional C++ (CC++) and Split-C, that provide support for a global name space. To establish a common comparison basis, our implementation of CC++ was developed to use MRPC, a RPC system optimized for MPMD parallel computing and based on Active Messages. Basic RPC performance in CC++ is within a factor of two from those of Split-C and other messaging layers. CC++ applications perform within a factor of two to six from comparable Split-C versions, which represent an order of magnitude improvement over previous CC++ implementations. The results suggest that RPC-based communication can be used effectively in many high-performance MPMD parallel applications. Chi-Chao Chang, Grzegorz Czajkowski, Thorsten von Eicken, Carl Kesselman |
SC | 4 |
| 1997 | Managing Multiple Communication Methods in High-Performance Networked Computing Systems
Ian T. Foster, Jonathan Geisler, Carl Kesselman, Steven Tuecke |
J. Parallel Distributed Comput. | 3 |
| 1996 | Multimethod Communication for High-Performance Metacomputing ApplicationsabstractMetacomputing systems use high-speed networks to connect supercomputers, mass storage systems, scientific instruments, and display devices with the objective of enabling parallel applications to utilize geographically distributed computing resources. However, experience shows that high performance can often be achieved only if applications can integrate diverse communication substrates, transport mechanisms, and protocols, chosen according to where communication is directed, what is communicated, or when communication is performed. In this paper, we describe a software architecture that addresses this requirement. This architecture allows multiple communication methods to be supported transparently in a single application, with either automatic or user-specified selection criteria guiding the methods used for each communication. We describe an implementation of this architecture, based on the Nexus communication library, and use this implementation to evaluate performance issues. This implementation was used to support a wide variety of applications in the I-WAY metacomputing experiment at Supercomputing~95; we use one of these applications to provide a quantitative demonstration of the advantages of multimethod communication in a heterogeneous networked environment. Ian T. Foster, Jonathan Geisler, Carl Kesselman, Steven Tuecke |
SC | 3 |
| 1996 | The Nexus Approach to Integrating Multithreading and Communication
Ian T. Foster, Carl Kesselman, Steven Tuecke |
J. Parallel Distributed Comput. | 2 |
| 1993 | Integrating task and data parallelismabstractNo abstract available. Ian T. Foster, Carl Kesselman |
SC | 2 |
| 1993 | Common runtime support for high-performance parallel languagesabstractNo abstract available. Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick |
SC | 15 |
| 1990 | Concurrency: Simple Concepts and Powerful ToolsabstractStepwise refinement is a central program development methodology that has been applied extensively to the design of sequential and parallel programs. In this methodology, a problem is successively decomposed into subproblems in order to untangle seemingly interdependent aspects of the design. To apply the methodology to parallel programs, one must be able to separate and reason about issues such as partitioning and mapping. This paper describes programming language concepts that we have found useful in applying stepwise refinement to parallel programs. The concepts allow decisions concerning program structure to be delayed until late in the design process. This capability permits rapid experimentation with alternative structures and leads to both portable and scalable code. Although simple, the concepts form a sufficient basis for the construction of powerful programming tools. Both concepts and tools have been applied successfully in a wide variety of applications and are incorporated in a commercial concurrent programming system, Strand*. Ian T. Foster, Carl Kesselman |
Comput. J. | 2 |