Kate Keahey

dblp:k/KatarzynaKeahey · also Katarzyna Keahey · DBLP profile ↗
← Back
41ranked-venue papers
19as first author
5since 2021 · last 2025
0000-0002-5251-5466ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 13 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 On the Reproducibility Challenges of Federated Learning: Investigating the Gap Between Simulation, Emulation and Real-World Deployments
abstract
Federated Learning (FL) is an emerging paradigm for decentralized training of Machine Learning models. It has been the subject of a large corpus of research due to its innovative approach to handling sensitive data. A common practice in the FL literature is to run simulations on a single compute node to assess the performance of FL algorithms. While simulation enables fast prototyping and validation of algorithmic concepts, it may face limitations in reproducing the real system's performance in heterogeneous environments such as the Computing Continuum, and particularly on resource-constrained Edge devices. Conversely, emulation on distributed testbeds offers more effective means to accurately reproduce the performance of real-world devices. However, to the best of our knowledge, no prior research has investigated the differences between simulation and emulation in FL experiments. In this paper, we study the complementarity of these approaches and discuss their respective challenges, as a first step towards reproducibility of FL experiments. We illustrate our study with a real-life application used as a baseline: an outdoor air quality forecasting framework with real-world sensors. Our results show that simulation can be used to accurately reproduce model performance metrics, while emulation can effectively reproduce the system performance of real-world experiments. Finally, we present a set of lessons learned on the challenges of FL reproducibility and the selection of experimental infrastructures for FL experiments and applications.
Cèdric Prigent, Kate Keahey, Alexandru Costan, Loïc Cudennec, Gabriel Antoniu
CCGrid2
2025 Lessons Learned from Anomaly Detection in Chameleon Cloud
abstract
Cloud computing has become integral to modern technology infrastructure, supporting a wide range of services from e-commerce to AI applications. Chameleon is a large-scale, configurable testbed designed to enable edge-to-cloud research through full bare-metal provisioning, virtualization, and diverse hardware resources, which is built on a leading open source cloud platform OpenStack. However, monitoring Chameleon’s heterogeneous infrastructure is challenging, particularly across Open-Stack services and hardware components. Traditional threshold-based alerting methods struggle to keep up with the scale and complexity of such environments. In this work, we present an anomaly detection framework for OpenStack services in the Chameleon Cloud. We curate and publish the first dataset of resource usage metrics collected from OpenStack control plane services. We evaluate four state-of-the-art unsupervised multivariate time series models, namely TranAD, Prodigy, USAD, and OmniAnomaly, on this dataset and share key insights from deploying them. Our findings indicate that for our use case, while all models achieve high F1 scores, training with three days of healthy data effectively balances training cost and detection accuracy.
S. M. Qasim, Can Hankendi, Kate Keahey, Gianluca Stringhini, Ayse K. Coskun
IC2E4
2025 Design and implementation of ARA wireless living lab for rural broadband and applications
Taimoor Ul Islam, Joshua Ofori Boateng, Md Nadim, Guoying Zu, Mukaram Shahid, Tianyi Zhang 0016, Salil Reddy, Wei Xu 0056, Ataberk Atalar, Vincent Lee, Yung-fu Chen, Evan Gossling, Elisabeth Permatasari, Christ Somiah, Owen Perrin, Zhibo Meng, Reshal Afzal, Sarath Babu 0001, Mohammed Soliman, Ali Hussain, Daji Qiao, Mai Zheng, Ozdal Boyraz, Anish Arora, Mohamed Y. Selim, Arsalan Ahmad, Myra B. Cohen, Mike Luby, Ranveer Chandra, James Gross, Kate Keahey, Hongwei Zhang 0001
Comput. Networks33
2023 Three Pillars of Practical Reproducibility
abstract
The ability to reproduce results is important for future experimenters building on existing work. We argue that the ability to do so in a cost-effective manner – what we call practical reproducibility – should be a mainstream method of scientific exploration and as a community we should invest in building an ecosystem of tools that would support it. Since computational research artifacts require some form of computing to interpret, we argue that such ecosystem is likely to be tied to shared experimental infrastructure where unique resources can be found and effectively shared. We propose three interrelated services, solving the problems of packaging for reuse, findability, and accessibility, respectively. We describe how we developed these services in Chameleon, an NSF-funded testbed for computer science research which has supported 8,000+ users, and discuss their strengths and limitations.
Kate Keahey, Mark Powers, Adam Cooper
e-Science1
2023 Discovery Testbed: An Observational Instrument for Broadband Research
abstract
Investigating phenomena that require continuous collection of data from a large and widely distributed array of hard-to-reach sources – such as understanding the performance of end-user broadband – has traditionally been hard. The opportunities in this sphere are now changing with easy availability of single board computers (SBCs), that are reliable and cheap, and thus can in principle be deployed in places of interest at large scales to gain coverage yielding statistically significant results. The challenge that this idea raises is how to create, deploy, an operate this type of infrastructure, given that its makup and deployment properties are very different from hardware deployed in datacenters. This paper presents the design of FLOTO, an observational instrument that supports the deployment and operation of mainstream SBCs to collect data through large-scale deployments in the field. FLOTO allows users to deploy devices to collect data of interest; operate those devices securely in remote locations without physical access to device; supports multi-tenant sharing of devices between different data collecting applications and user groups, that makes it possible to adapt or re-purpose the observational function of the instrument; and provides data collection and sharing functionality allowing user communities to benefit from the collected data. We describe the design of FLOTO and present a case study of its deployment as an observational instrument to collect broadband data. We conclude by discussing the possible adaptations of this instrument to study different scientific questions.
Kate Keahey, Nick Feamster, Guilherme Martins, Mark Powers, Marc Richardson, Alexis Schrubbe
e-Science1
2020 Lessons Learned from the Chameleon Testbed
Kate Keahey, Zhuo Zhen, Pierre Riteau, Paul Ruth, Daniel C. Stanzione Jr., Mert Cevik, Jacob Colleran, Haryadi S. Gunawi, Cody Hammock, Joe Mambretti, Alexander Barnes, François Halbach, Alex Rocha, Joe Stubbs
USENIX ATC1
2019 Managing Allocatable Resources
abstract
Infrastructure cloud computing allows its clients to allocate on-demand resources, typically consisting of a representation of a compute node. In general however, there is a need for allocating resources other than nodes and managing them in more controlled ways than simply on demand. This paper generalizes the familiar "compute power on demand" pattern by introducing the abstraction of an allocatable resource, describing its properties, and implementation for different types of resources. We further describe architecture for a generic allocatable resource management service that can be extended to manage diverse types of resources as well as the implementation of this architecture in the OpenStack Blazar service to manage resources ranging from bare-metal compute nodes to network segments. Finally, we provide a usage analysis of this service on the Chameleon testbed and use it to illustrate the effectiveness of resource management methods as well as the need for incentives in usage arbitration.
Kate Keahey, Pierre Riteau, Zhuo Zhen
CLOUD1
2019 A Case for Integrating Experimental Containers with Notebooks
abstract
Computational notebooks have gained much popularity as a way of documenting research processes; they allow users to express research narrative by integrating ideas expressed as text, process expressed as code, and results in one executable document. However, the environments in which the code can run are currently limited, often containing only a fraction of the resources of one node, posing a barrier to many computations. In this paper, we make the case that integrating complex experimental environments, such as virtual clusters or complex networking environments that can be provisioned via infrastructure clouds, into computational notebooks will significantly broaden their reach and at the same time help realize the potential of clouds as a platform for repeatable research. To support our argument, we describe the integration of Jupyter notebooks into the Chameleon cloud testbed, which allows the user to define complex experimental environments and then assign processes to elements of this environment similarly to the way a laptop user may switch between different desktops. We evaluate our approach on an actual experiment from both the development and replication perspective.
Kate Keahey
CloudCom2
2019 Chameleon: A Large-Scale, Deeply Reconfigurable Testbed for Computer Science Research
abstract
Computer Science experimental testbeds allow investigators to explore a broad range of different state-of-the-art hardware options, assess scalability of their systems, and provide conditions that allow deep reconfigurability and isolation so that one user does not impact the experiments of another. An experimental testbed is also in a unique position to support methods facilitating experiment analysis and improve repeatability and reproducibility of experiments. Providing these capabilities at least partially within a commodity framework improves the sustainability of systems experiments and thus makes them available to a broader range of experimenters.
Kate Keahey, Joe Mambretti, Paul Ruth, Daniel C. Stanzione Jr.
ICNP1
2019 Reducing Kernel Surface Areas for Isolation and Scalability
abstract
Isolation is a desirable property for applications executing in multi-tenant computing systems. On the performance side, hardware resource isolation via partitioning mechanisms is commonly applied to achieve QoS, a necessary property for many noise-sensitive parallel workloads. Conversely, on the software side, partitioning is used, usually in the form of virtual machines, to provide secure environments with smaller attack surfaces than those present in shared software stacks.
Daniel Zahka, Brian Kocoloski, Kate Keahey
ICPP3
2018 Dynamically negotiating capacity between on-demand and batch clusters
Kate Keahey, Pierre Riteau, Jon B. Weissman
SC2
2017 Traffic-sensitive Live Migration of Virtual Machines
Umesh Deshpande, Kate Keahey
Future Gener. Comput. Syst.2
2016 Guest Editors Introduction: Special Issue on Scientific Cloud Computing
abstract
The papers in this special section contribute important advances towards leveraging clouds for scientific applications. The contributions focus on a broad range of topics, including: performance modeling and optimization, data management, resource allocation and scheduling, elasticity, reconfiguration, cost prediction and optimization. Most papers revolve around general techniques and approaches that are agnostic of the applications, while two contributions demonstrate how domain specific scientific applications can be migrated to the cloud.
Kate Keahey, Ioan Raicu, Kyle Chard, Bogdan Nicolae
IEEE Trans. Cloud Comput.1
2015 Traffic-Sensitive Live Migration of Virtual Machines
abstract
In this paper we address the problem of network contention between the migration traffic and the VM application traffic for the live migration of co-located Virtual Machines (VMs). When VMs are migrated with pre-copy, they run at the source host during the migration. Therefore the VM applications with predominantly outbound traffic contend with the outgoing migration traffic at the source host. Similarly, during post-copy migration, the VMs run at the destination host. Therefore the VM applications with predominantly inbound traffic contend with the incoming migration traffic at the destination host. Such a contention increases the total migration time of the VMs and degrades the performance of VM application. Here, we propose traffic-sensitive live VM migration technique to reduce the contention of migration traffic with the VM application traffic. It uses a combination of pre-copy and post-copy techniques for the migration of the co-located VMs, instead of relying upon any single pre-determined technique for the migration of all the VMs. We base the selection of migration techniques on VMs' network traffic profiles so that the direction of migration traffic complements the direction of the most VM application traffic. We have implemented a prototype of traffic-sensitive migration on the KVM/QEMU platform. In the evaluation, we compare traffic-sensitive migration against the approaches that use only pre-copy or only post-copy for VM migration. We show that our approach minimizes the network contention for migration, thus reducing the total migration time and the application degradation.
Umesh Deshpande, Kate Keahey
CCGRID2
2014 Evaluating Streaming Strategies for Event Processing Across Infrastructure Clouds
abstract
Infrastructure clouds revolutionized the way in which we approach resource procurement by providing an easy way to lease compute and storage resources on short notice, for a short amount of time, and on a pay-as-you-go basis. This new opportunity, however, introduces new performance trade-offs. Making the right choices in leveraging different types of storage available in the cloud is particularly important for applications that depend on managing large amounts of data within and across clouds. An increasing number of such applications conform to a pattern in which data processing relies on streaming the data to a compute platform where a set of similar operations is repeatedly applied to independent chunks of data. This pattern is evident in virtual observatories such as the Ocean Observatory Initiative, in cases when new data is evaluated against existing features in geospatial computations or when experimental data is processed as a series of time events. In this paper, we propose two strategies for efficiently implementing such streaming in the cloud and evaluate them in the context of an ATLAS application processing experimental data. Our results show that choosing the right cloud configuration can improve overall application performance by as much as three times.
Radu Tudoran, Kate Keahey, Pierre Riteau, Sergey Panitkin, Gabriel Antoniu
CCGRID2
2014 Bursting the Cloud Data Bubble: Towards Transparent Storage Elasticity in IaaS Clouds
abstract
Storage elasticity on IaaS clouds is an important feature for data-intensive workloads: storage requirements can vary greatly during application runtime, making worst-case over-provisioning a poor choice that leads to unnecessarily tied-up storage and extra costs for the user. While the ability to adapt dynamically to storage requirements is thus attractive, how to implement it is not well understood. Current approaches simply rely on users to attach and detach virtual disks to the virtual machine (VM) instances and then manage them manually, thus greatly increasing application complexity while reducing cost efficiency. Unlike such approaches, this paper aims to provide a transparent solution that presents a unified storage space to the VM in the form of a regular POSIX file system that hides the details of attaching and detaching virtual disks by handling those actions transparently based on dynamic application requirements. The main difficulty in this context is to understand the intent of the application and regulate the available storage in order to avoid running out of space while minimizing the performance overhead of doing so. To this end, we propose a storage space prediction scheme that analyzes multiple system parameters and dynamically adapts monitoring based on the intensity of the I/O in order to get as close as possible to the real usage. We show the value of our proposal over static worst-case over-provisioning and simpler elastic schemes that rely on a reactive model to attach and detach virtual disks, using both synthetic benchmarks and real-life data-intensive applications. Our experiments demonstrate that we can reduce storage waste/cost by 30-40% with only 2-5% performance overhead.
Bogdan Nicolae, Pierre Riteau, Kate Keahey
IPDPS3
2012 Architecting a Large-scale Elastic Environment - Recontextualization and Adaptive Cloud Services for Scientific Computing
abstract
Infrastructure-as-a-service (IaaS) clouds, such as Amazon EC2, offer pay-for-use virtual resources ondemand. This allows users to outsource computation and storage when needed and create elastic computing environments that adapt to changing demand. However, existing services, such as cluster resource managers (e.g. Torque), do not include support for elastic environments. Furthermore, no recontextualization services exist to reconfigure these environments as they continually adapt to changes in demand. In this paper we present an architecture for a large-scale elastic cluster environment. We extend an open-source elastic IaaS manager, the Elastic Processing Unit (EPU), to support the Torque batch-queue scheduler. We also develop a lightweight REST-based recontextualization broker that periodically reconfigures the cluster as nodes join or leave the environment. Our solution adds nodes dynamically at runtime and supports MPI jobs across distributed resources. For experimental evaluation, we deploy our solution using both NSF FutureGrid and Amazon EC2. We demonstrate the ability of our solution to create multi-cloud deployments and run batchqueued jobs, recontextualize 256 node clusters within one second of the recontextualization period, and scale to over 475 nodes in less than 15 minutes.
Paul Marshall, Henry M. Tufo, Kate Keahey, David La Bissoniere, Matthew Woitaszek
ICSOFT3
2011 Improving Utilization of Infrastructure Clouds
abstract
A key advantage of infrastructure-as-a-service (IaaS) clouds is providing users on-demand access to resources. To provide on-demand access, however, cloud providers must either significantly overprovision their infrastructure (and pay a high price for operating resources with low utilization) or reject a large proportion of user requests (in which case the access is no longer on-demand). At the same time, not all users require truly on-demand access to resources. Many applications and workflows are designed for recoverable systems where interruptions in service are expected. For instance, many scientists utilize high-throughput computing (HTC)-enabled resources, such as Condor, where jobs are dispatched to available resources and terminated when the resource is no longer available. We propose a cloud infrastructure that combines on-demand allocation of resources with opportunistic provisioning of cycles from idle cloud nodes to other processes by deploying backfill virtual machines (VMs). For demonstration and experimental evaluation, we extend the Nimbus cloud computing toolkit to deploy backfill VMs on idle cloud nodes for processing an HTC workload. Initial tests show an increase in IaaS cloud utilization from 37.5% to 100% during a portion of the evaluation trace but only 6.39% overhead cost for processing the HTC workload. We demonstrate that a shared infrastructure between IaaS cloud providers and an HTC job management system can be highly beneficial to both the IaaS cloud provider and HTC users by increasing the utilization of the cloud infrastructure (thereby decreasing the overall cost) and contributing cycles that would otherwise be idle to processing HTC jobs.
Paul Marshall, Kate Keahey, Timothy Freeman 0001
CCGRID2
2011 Going back and forth: efficient multideployment and multisnapshotting on clouds
abstract
Infrastructure as a Service (IaaS) cloud computing has revolutionized the way we think of acquiring resources by introducing a simple change: allowing users to lease computational resources from the cloud provider's datacenter for a short time by deploying virtual machines (VMs) on these resources. This new model raises new challenges in the design and development of IaaS middleware. One of those challenges is the need to deploy a large number (hundreds or even thousands) of VM instances simultaneously. Once the VM instances are deployed, another challenge is to simultaneously take a snapshot of many images and transfer them to persistent storage to support management tasks, such as suspend-resume and migration. With datacenters growing rapidly and configurations becoming heterogeneous, it is important to enable efficient concurrent deployment and snapshotting that are at the same time hypervisor independent and ensure a maximum compatibility with different configurations. This paper addresses these challenges by proposing a virtual file system specifically optimized for virtual machine image storage. It is based on a lazy transfer scheme coupled with object versioning that handles snapshotting transparently in a hypervisor-independent fashion, ensuring high portability for different configurations. Large-scale experiments on hundreds of nodes demonstrate excellent performance results: speedup for concurrent VM deployments ranges from a factor of 2 up to 25, with a reduction in bandwidth utilization of as much as 90%.
Bogdan Nicolae, John Bresnahan, Kate Keahey, Gabriel Antoniu
HPDC3
2010 Elastic Site: Using Clouds to Elastically Extend Site Resources
abstract
Infrastructure-as-a-Service (IaaS) cloud computing offers new possibilities to scientific communities. One of the most significant is the ability to elastically provision and relinquish new resources in response to changes in demand. In our work, we develop a model of an “elastic site” that efficiently adapts services provided within a site, such as batch schedulers, storage archives, or Web services to take advantage of elastically provisioned resources. We describe the system architecture along with the issues involved with elastic provisioning, such as security, privacy, and various logistical considerations. To avoid over- or under-provisioning the resources we propose three different policies to efficiently schedule resource deployment based on demand. We have implemented a resource manager, built on the Nimbus toolkit to dynamically and securely extend existing physical clusters into the cloud. Our elastic site manager interfaces directly with local resource managers, such as Torque. We have developed and evaluated policies for resource provisioning on a Nimbus-based cloud at the University of Chicago, another at Indiana University, and Amazon EC2. We demonstrate a dynamic and responsive elastic cluster, capable of responding effectively to a variety of job submission patterns. We also demonstrate that we can process 10 times faster by expanding our cluster up to 150 EC2 nodes.
Paul Marshall, Kate Keahey, Timothy Freeman 0001
CCGRID2
2010 Grid, Cluster and Cloud Computing
Kate Keahey, Domenico Laforenza, Alexander Reinefeld, Pierluigi Ritrovato, Douglas Thain, Nancy Wilkins-Diehr
Euro-Par (1)1
2009 Cloud Computing for Science
Kate Keahey
SSDBM1
2008 On the Use of Cloud Computing for Scientific Workflows
abstract
This paper explores the use of cloud computing for scientific workflows, focusing on a widely used astronomy application-Montage. The approach is to evaluate from the point of view of a scientific workflow the tradeoffs between running in a local environment, if such is available, and running in a virtual environment via remote, wide-area network resource access. Our results show that for Montage, a workflow with short job runtimes, the virtual environment can provide good compute time performance but it can suffer from resource scheduling delays and widearea communications.
Christina Hoffa, Gaurang Mehta, Timothy Freeman 0001, Ewa Deelman, Kate Keahey, G. Bruce Berriman, John Good
eScience5
2008 Contextualization: Providing One-Click Virtual Clusters
abstract
As virtual appliances become more prevalent, we encounter the need to stop manually adapting them to their deployment context each time they are deployed. We examine appliance contextualization needs and present architecture for secure, consistent, and dynamic contextualization, in particular for groups of appliances that must work together in a shared security context. This architecture allows for programmatic cluster creation and use, as well as mitigating potential errors and unnecessary charges during setup time. For portability across many deployment mechanisms, we introduce the concept of a standalone context broker. We describe the current implementation of the entire architecture using the virtual workspaces toolkit, showing real-life examples of dynamically contextualized Grid clusters.
Kate Keahey, Timothy Freeman 0001
eScience1
2008 Flying Low: Simple Leases with Workspace Pilot
Timothy Freeman 0001, Kate Keahey
Euro-Par2
2008 Combining batch execution and leasing using virtual machines
abstract
As cluster computers are used for a wider range of applications, we encounter the need to deliver resources at particular times, to meet particular deadlines, and/or at the same time as other resources are provided elsewhere. To address such requirements, we describe a scheduling approach in which users request resource leases, where leases can request either as-soon-as-possible ("best-effort") or reservation start times. We present the design of a lease management architecture, Haizea, that implements leases as virtual machines (VMs), leveraging their ability to suspend, migrate, and resume computations and to provide leased resources with customized application environments. We discuss methods to minimize the overhead introduced by having to deploy VM images before the start of a lease. We also present the results of simulation studies that compare alternative approaches. Using workloads with various mixes of best-effort and advance reservation requests, we compare the performance of our VM-based approach with that of non-VM-based schedulers. We find that a VM-based approach can provide better performance (measured in terms of both total execution time and average delay incurred by best-effort requests) than a scheduler that does not support task pre-emption, and only slightly worse performance than a scheduler that does support task pre-emption. We also compare the impact of different VM image popularity distributions and VM image caching strategies on performance. These results emphasize the importance of VM image caching for the workloads studied and quantify the sensitivity of scheduling performance to VM image popularity distribution.
Borja Sotomayor, Kate Keahey, Ian T. Foster
HPDC2
2006 Virtual Clusters for Grid Communities
abstract
A challenging issue facing Grid communities is that while Grids can provide access to many heterogeneous resources, the resources to which access is provided often do not match the needs of a specific application or service. In an environment in which both resource availability and software requirements evolve rapidly, this disconnect can lead to resource underutilization, user frustration, and much wasted effort spent on bridging the gap between applications and resources. We show here how these issues can be overcome by allowing authorized Grid clients to negotiate the creation of virtual clusters made up of virtual machines configured to suit client requirements for software environment and hardware allocation. We introduce descriptions and methods that allow us to deploy flexibly configured virtual cluster workspaces. We describe their configuration, implementation, and evaluate them in the context of a virtual cluster representing the environment in production use by the Open Science Grid. Our performance evaluation results show that virtual clusters representing current Grid production environments can be deployed and managed efficiently, and thus can provide an acceptable platform for Grid applications.
Ian T. Foster, Timothy Freeman 0001, Kate Keahey, Doug Scheftner, Borja Sotomayor, Xuehai Zhang
CCGRID3
2006 Division of Labor: Tools for Growing and Scaling Grids
Timothy Freeman 0001, Kate Keahey, Ian T. Foster, Abhishek Singh Rana, Borja Sotomayor, Frank Würthwein
ICSOC2
2006 Virtual playgrounds: managing virtual resources in the grid
abstract
Large grid deployments increasingly require abstractions and methods decoupling the work of resource providers and resource consumers to implement scalable management methods. We proposed the abstraction of a virtual workspace (VW) describing a virtual execution environment that can be made dynamically available to authorized grid clients by using well-defined protocols. Virtual workspaces provide resources in controllable ways that are independent of how a resource is consumed. A virtual playground may combine many such workspaces, as well as other aspects of virtual environments, such as networking and storage, to form virtual grids. In this paper, we report on the goals and progress of the virtual playground project and put in context the research to date.
Kate Keahey, Jeffrey S. Chase, Ian T. Foster
IPDPS1
2006 Virtualization technologies - 1st IEEE/ACM international workshop on virtualization technologies in distributed computing
abstract
The convergence of virtualization technologies and distributed computing is an exciting ongoing development and the subject of much current research. This workshop is intended as a forum for the exchange of ideas and experiences on the use of virtualization technologies in distributed computing, the challenges and opportunities offered by the development of virtual systems themselves, and case studies of application of virtualization. Virtualization technologies include virtual machines, virtual networks, virtual data, virtual storage, virtual applications, and virtual instruments. Distributed computing includes Grid computing, cluster computing, peer-to-peer computing and mobile computing.
Kate Keahey
SC1
2006 Poster reception - To bid or not to bid: a hybrid market-based resource allocation framework
abstract
In this work we present a novel market based resource allocation framework named Virtual Workspaces Market. Our approach is to integrate Virtual Workspaces and some Tycoon components while developing while developing a new hybrid resource allocation policy which combines the advantages of leasing and auctions based mechanisms. The rationale behind the hybrid allocation policy is to have an efficient and flexible resource sharing environment that fulfill the conflicting user/provider requirements by exploiting the clients demands heterogeneity. Our solution builds on current advances of Virtualization and Grid technologies (i.e. an integration and extension of Virtual Workspaces, Tycoon and Xen Hypervisor components). The main contributions are: (I) A market-based framework that combines advantages of leasing and auction resource allocation models. (II) An analysis of a hybrid market based resource allocation policy, where users can choose between paying more to have a low-risk resource reservation or paying less and controlling resource shares via bids.
Elizeu Santos-Neto, Kate Keahey
SC2
2005 Virtual Workspaces in the Grid
Kate Keahey, Ian T. Foster, Timothy Freeman 0001, Xuehai Zhang, Daniel Galron
Euro-Par1
2004 Agreement-Based Interactions for Experimental Science
Kate Keahey, Takuya Araki, Peter Lane
Euro-Par1
2004 Fine-grained authorization for job execution in the Grid: design and implementation
abstract
Abstract In this paper, we describe our work on enabling fine‐grained authorization for resource usage and management. We address the need of virtual organizations to enforce their own polices in addition to those of the resource owners, with regards to both resource consumption and job management. To implement this design, we propose changes and extensions to the Globus Toolkit's version 2 resource management mechanism. We describe the prototype and policy language that we have designed to express fine‐grained policies and present an analysis of our solution. Copyright © 2004 John Wiley & Sons, Ltd.
Kate Keahey, Von Welch, Sam Lang, Sam Meder
Concurr. Pract. Exp.1
2002 A CORBA Commodity Grid Kit
abstract
Abstract This paper reports on an ongoing research project aimed at designing and deploying a Common Object Resource Broker Architecture (CORBA) (ww.omg.org) Commodity Grid (CoG) Kit. The overall goal of this project is to enable the development of advanced Grid applications while adhering to state‐of‐the‐art software engineering practices and reusing the existing Grid infrastructure. As part of this activity, we are investigating how CORBA can be used to support the development of Grid applications. In this paper, we outline the design of a CORBA CoG Kit that will provide a software development framework for building a CORBA ‘Grid domain’. We also present our experiences in developing a prototype CORBA CoG Kit that supports the development and deployment of CORBA applications on the Grid by providing them access to the Grid services provided by the Globus Toolkit. Copyright © 2002 John Wiley & Sons, Ltd.
Manish Parashar, Gregor von Laszewski, Snigdha Verma, Jarek Gawor, Kate Keahey, Nell Rehn
Concurr. Comput. Pract. Exp.5
2002 Computational Grids in action: the National Fusion Collaboratory
Kate Keahey, Thomas W. Fredian, David P. Schissel, Mary R. Thompson, Ian T. Foster, M. J. Greenwald, Douglas McCune
Future Gener. Comput. Syst.1
2001 PAWS: Collective Interactions and Data Transfers
abstract
The authors discuss problems and solutions pertaining to the interaction of components representing parallel applications. We introduce the notion of a collective port which is an extension of the Common Component Architecture (CCA) ports and allows collective components representing parallel applications to interact as one entity. We further describe a class of translation components, which translate between the distributed data format used by one parallel implementation to that used by another. A well known example of such components is the MxN component which translates between data distributed on M processors to data distributed on N processors. We describe its implementation in Parallel Application Work Space (PAWS), as well as the data structures PAWS uses to support it. We also present a mechanism allowing the framework to invoke this component on the programmer's behalf whenever such translation is necessary, freeing the programmer from treating collective component interactions as a special case. In doing that, we introduce framework-based, user-defined distributed type casts. Finally, we discuss our initial experiments in building optimized complex translation components out of atomic functionalities.
Kate Keahey, Patricia K. Fasel, Susan M. Mniszewski
HPDC1
1999 Toward a Common Component Architecture for High-Performance Scientific Computing
abstract
Describes work in progress to develop a standard for interoperability among high-performance scientific components. This research stems from the growing recognition that the scientific community needs to better manage the complexity of multidisciplinary simulations and better address scalable performance issues on parallel and distributed architectures. The driving force for this is the need for fast connections among components that perform numerically intensive work and for parallel collective interactions among components that use multiple processes or threads. This paper focuses on the areas we believe are most crucial in this context, namely an interface definition language that supports scientific abstractions for specifying component interfaces and a port connection model for specifying component interactions.
Robert C. Armstrong, Dennis Gannon, Al Geist, Kate Keahey, Scott R. Kohn, Lois C. McInnes, Steven G. Parker, Brent A. Smolinski
HPDC4
1999 PARDIS: Programmer-level abstractions for metacomputing
Kate Keahey
Future Gener. Comput. Syst.1
1997 PARDIS: A Parallel Approach to CORBA
abstract
This paper describes PARDIS, a system containing explicit support for interoperability of PARallel DIStributed applications. PARDIS is based on the Common Object Request Broker Architecture (CORBA). Like CORBA, it provides interoperability between heterogeneous components by specifying their interfaces in a meta-language, the CORBA IDL, which call be translated into the language of interacting components. However, PARDIS extends the CORBA object model by introducing SPMD objects representing data-parallel computations. SPMD objects allow the request broker to interact directly with the distributed resources of a parallel application. This capability ensures request delivery to all the computing threads of a parallel application and allows the request broker to transfer distributed arguments directly between the computing threads of the client and the server. To support this kind of argument transfer, PARDIS defines a distributed argument type-distributed sequence-a generalization of CORBA sequence representing distributed data structures of parallel applications. In this paper we will give a brief description of basic component interaction in PARDIS and give an account of the rationale and support for SPMD objects and distributed sequences. We will then describe two ways of implementing argument transfer in invocations on SPMD objects and evaluate and compare their performance.
Kate Keahey, Dennis Gannon
HPDC1
1997 PARDIS: CORBA-based Architecture for Application-Level Parallel Distributed Computation
abstract
Modern technology provides the infrastructure necessary to develop distributed applications capable of using the power of multiple supercomputing resources and exploiting their diversity. The performance potential offered by distributed supercomputing is enormous, but it is hard to realize due to the complexity of programming in such environments. In this paper we introduce PARDIS, a system designed to overcome this challenge, based on ideas underlying the Common Object Request Broker Architecture (CORBA), a successful industry standard. PARDIS is a distributed environment in which objects representing data-parallel computations, called SPMD objects, as well as non-parallel objects present in parallel programs, can interact with each other across platforms and software systems. Each of these objects represents a small encapsulated application and can be used as a building block in the construction of powerful distributed metaapplications. The objects interact through interfaces specified in the Interface Definition Language (IDL), which allows the programmer to integrate within one metaapplication components implemented using different software systems. Further, support for non-blocking interactions between objects allows PARDIS to build concurrent distributed scenarios.
Kate Keahey, Dennis Gannon
SC1