Jeffrey S. Chase

dblp:c/JSChase · DBLP profile ↗
← Back
60ranked-venue papers
7as first author
5since 2021 · last 2022
0000-0001-8275-0830ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 18 · 4 first-authorComputer networks · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 NetHint: White-Box Networking for Multi-Tenant Data Centers
Jingrong Chen 0002, Hong Zhang 0025, Jeffrey S. Chase, Ion Stoica, Danyang Zhuo
NSDI5
2022 ImPACT: A networked service architecture for safe sharing of restricted data
abstract
In this paper we describe an architecture developed and prototyped in the course of the NSF-funded project called ImPACT—Infrastructure for Privacy-Assured CompuTations. This architecture addresses the common problems that arise from the need to securely store, control access to and process privacy-restricted data in a multi-institutional, multi-stakeholder setting. Specifically the architecture includes several components—a way to publicly advertise a limited set of data attributes without exposing the sensitive data itself; a set of mechanisms for a data owner to specify and automatically enforce complex data-access policies commonly expressed today as Data Use Agreements (DUAs); a way to securely collect digital attestations from multiple stakeholders to satisfy those policies; and a reproducible template to deploy secure processing enclaves in which groups of researchers can analyze the data in a way that complies with data owner policies using the tools of their choice. The paper describes the architecture and its instantiation in a prototype, providing a performance evaluation of several components.
Ilya Baldin, Jeffrey S. Chase, Jonathan David Crabtree, Thomas Nechyba, Laura Christopherson, Michael J. Stealey, Charles Kneifel, Victor Orlikowski, Rob Carter, Erik Scott, Akio Sone, Donald Sizemore
Future Gener. Comput. Syst.2
2021 WIRE: Resource-efficient Scaling with Online Prediction for DAG-based Workflows
abstract
This paper introduces WIRE that manages resources for the DAG-based workflows on IaaS clouds. WIRE predicts and plans resources over the MAPE (Monitor-Analyze-Plan-Execute) loops to: 1) Estimate task performance with online data, 2) Conduct simulations to predict the upcoming loads based on online estimates and workflow DAGs, 3) Apply a resource-steering policy to size cloud instance pools for the maximal parallelism that is consistent with low cost. We implement WIRE on Pegasus WMS/HTCondor and evaluate its performance on the ExoGENI network cloud. The results show that WIRE attains low resource cost with the performance that is typically within a factor of two of optimal.
Qiang Cao 0005, Mayuresh Kunjir, Linli Wan, Jeffrey S. Chase, Anirban Mandal, Mats Rynge
CLUSTER5
2021 Federated Authorization for Managed Data Sharing: Experiences from the ImPACT Project
abstract
This paper presents the rationale and design of the trust plane for ImPACT, a federated platform for managed sharing of restricted data. Key elements of the architecture include Web-based notaries for credential establishment based on declarative templates for Data Usage Agreements, a federated authorization pipeline, integration of popular services for identity management, and programmable policy based on a logical trust model with a repository of linked certificates. We show how these elements of the trust plane work in concert, and set the ideas in context with principles of federated authorization. A focus and contribution of the paper is to explore limitations of the resulting architecture and tensions among competing design goals. We also point the way toward future extensions, including policy-checked data access from cloud-hosted data enclaves with enhanced defenses against data leakage and exfiltration.
Jeffrey S. Chase, Ilya Baldin
ICCCN1
2021 Interpreting Write Performance of Supercomputer I/O Systems with Regression Models
abstract
This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10 x on some samples for both of the target systems.
Zilong Tan, Philip H. Carns, Jeffrey S. Chase, Kevin Harms, Jay F. Lofstead, Sarp Oral, Sudharshan S. Vazhkudai, Feiyi Wang
IPDPS4
2020 Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and Implications
abstract
This article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries. We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications. In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable.
Sarp Oral, Christopher Zimmer 0001, Jong Choi 0001, David Dillow, Scott Klasky, Jay F. Lofstead, Norbert Podhorszki, Jeffrey S. Chase
ACM Trans. Storage9
2017 Enabling Lightweight Transactions with Precision Time
abstract
Distributed transactional storage is an important service in today's data centers. Achieving high performance without high complexity is often a challenge for these systems due to sophisticated consistency protocols and multiple layers of abstraction. In this paper we show how to combine two emerging technologies---Software-Defined Flash (SDF) and precise synchronized clocks---to improve performance and reduce complexity for transactional storage within the data center.
Pulkit A. Misra, Jeffrey S. Chase, Johannes Gehrke, Alvin R. Lebeck
ASPLOS2
2017 Predicting Output Performance of a Petascale Supercomputer
abstract
In this paper, we develop a predictive model useful for output performance prediction of supercomputer file systems under production load. Our target environment is Titan---the 3rd fastest supercomputer in the world---and its Lustre-based multi-stage write path. We observe from Titan that although output performance is highly variable at small time scales, the mean performance is stable and consistent over typical application run times. Moreover, we find that output performance is non-linearly related to its correlated parameters due to interference and saturation on individual stages on the path. These observations enable us to build a predictive model of expected write times of output patterns and I/O configurations, using feature transformations to capture non-linear relationships. We identify the candidate features based on the structure of the Lustre/Titan write path, and use feature transformation functions to produce a model space with 135,000 candidate models. By searching for the minimal mean square error in this space we identify a good model and show that it is effective.
Yezhou Huang, Jeffrey S. Chase, Jong Choi 0001, Scott Klasky, Jay F. Lofstead, Sarp Oral
HPDC3
2016 CQSTR: Securing Cross-Tenant Applications with Cloud Containers
abstract
Cloud providers are in a position to greatly improve the trust clients have in network services: IaaS platforms can isolate services so they cannot leak data, and can help verify that they are securely deployed. We describe a new system called CQSTR that allows clients to verify a service's security properties. CQSTR provides a new cloud container abstraction similar to Linux containers but for VM clusters within IaaS clouds. Cloud containers enforce constraints on what software can run, and control where and how much data can be communicated across service boundaries. With CQSTR, IaaS providers can make assertions about the security properties of a service running in the cloud.
Yan Zhai, Lichao Yin, Jeffrey S. Chase, Thomas Ristenpart, Michael M. Swift
SoCC3
2014 GENI: A federated testbed for innovative network experiments
Mark Berman 0001, Jeffrey S. Chase, Lawrence H. Landweber, Akihiro Nakao, Maximilian Ott, Dipankar Raychaudhuri, Robert Ricci, Ivan Seskar
Comput. Networks2
2012 Dynamic network provisioning for data intensive applications in the cloud
abstract
Advanced networks are an essential element of data-driven science enabled by next generation cyberinfrastructure environments. Computational activities increasingly incorporate widely dispersed resources with linkages among software components spanning multiple sites and administrative domains. We have seen recent advances in enabling on-demand network circuits in the national and international backbones coupled with Software Defined Networking (SDN) advances like OpenFlow and programmable edge technologies like OpenStack. These advances have created an unprecedented opportunity to enable complex scientific applications to run on specially tailored, dynamic infrastructure that include compute, storage and network resources, combining the performance advantages of purpose-built infrastructures, but without the costs of a permanent infrastructure. This work presents an experience deploying scientific workflows on the ExoGENI national test bed that dynamically allocates computational resources with high-speed circuits from backbone providers. Dynamically allocated bandwidth-provisioned high-speed circuits increase the ability of scientific applications to access and stage large data sets from remote data repositories or to move computation to remote sites and access data stored locally. The remainder of this extended abstract is a brief description of the test bed and several scientific workflow applications that were deployed using bandwidth-provisioned high-speed circuits.
Paul Ruth, Anirban Mandal, Yufeng Xin, Ilya Baldin, Chris Heermann, Jeffrey S. Chase
eScience6
2012 Characterizing output bottlenecks in a supercomputer
abstract
Supercomputer I/O loads are often dominated by writes. HPC (High Performance Computing) file systems are designed to absorb these bursty outputs at high bandwidth through massive parallelism. However, the delivered write bandwidth often falls well below the peak. This paper characterizes the data absorption behavior of a center-wide shared Lustre parallel file system on the Jaguar supercomputer. We use a statistical methodology to address the challenges of accurately measuring a shared machine under production load and to obtain the distribution of bandwidth across samples of compute nodes, storage targets, and time intervals. We observe and quantify limitations from competing traffic, contention on storage servers and I/O routers, concurrency limitations in the client compute node operating systems, and the impact of variance (stragglers) on coupled output such as striping. We then examine the implications of our results for application performance and the design of I/O middleware systems on shared supercomputers.
Jeffrey S. Chase, David Dillow, Oleg Drokin, Scott Klasky, Sarp Oral, Norbert Podhorszki
SC2
2011 Provisioning and Evaluating Multi-domain Networked Clouds for Hadoop-based Applications
abstract
This paper presents the design, implementation, and evaluation of a new system for on-demand provisioning of Hadoop clusters across multiple cloud domains. The Hadoop clusters are created "on-demand" and are composed of virtual machines from multiple cloud sites linked with bandwidth-provisioned network pipes. The prototype uses an existing federated cloud control framework called Open Resource Control Architecture (ORCA), which orchestrates the leasing and configuration of virtual infrastructure from multiple autonomous cloud sites and network providers. ORCA enables computational and network resources from multiple clouds and network substrates to be aggregated into a single virtual "slice" of resources, built to order for the needs of the application. The experiments examine various provisioning alternatives by evaluating the performance of representative Hadoop benchmarks and applications on resource topologies with varying bandwidths. The evaluations examine conditions in which multi-cloud Hadoop deployments pose significant advantages or disadvantages during Map/Reduce/Shuffle operations. Further, the experiments compare multi-cloud Hadoop deployments with single-cloud deployments and investigate Hadoop Distributed File System (HDFS) performance under varying network configurations. The results show that networked clouds make cross-cloud Hadoop deployment feasible with high bandwidth network links between clouds. As expected, performance for some benchmarks degrades rapidly with constrained inter-cloud bandwidth. MapReduce shuffle patterns and certain Hadoop Distributed File System (HDFS) operations that span the constrained links are particularly sensitive to network performance. Hadoop's topology-awareness feature can mitigate these penalties to a modest degree in these hybrid bandwidth scenarios. Additional observations show that contention among co-located virtual machines is a source of irregular performance for Hadoop applications on virtual cloud infrastructure.
Anirban Mandal, Yufeng Xin, Ilya Baldin, Paul Ruth, Chris Heermann, Jeffrey S. Chase, Victor Orlikowski, Aydan R. Yumerefendi
CloudCom6
2011 Deadline-sensitive workflow orchestration without explicit resource control
Lavanya Ramakrishnan, Jeffrey S. Chase, Dennis Gannon, Daniel Nurmi, Richard Wolski
J. Parallel Distributed Comput.2
2009 Rethinking FTP: Aggressive block reordering for large file transfers
abstract
Whole-file transfer is a basic primitive for Internet content dissemination. Content servers are increasingly limited by disk arm movement, given the rapid growth in disk density, disk transfer rates, server network bandwidth, and content size. Individual file transfers are sequential, but the block access sequence on a content server is effectively random when many slow clients access large files concurrently. Although larger blocks can help improve disk throughput, buffering requirements increase linearly with block size. This article explores a novel block reordering technique that can reduce server disk traffic significantly when large content files are shared. The idea is to transfer blocks to each client in any order that is convenient for the server. The server sends blocks to each client opportunistically in order to maximize the advantage from the disk reads it issues to serve other clients accessing the same file. We first illustrate the motivation and potential impact of aggressive block reordering using simple analytical models. Then we describe a file transfer system using a simple block reordering algorithm, called Circus. Experimental results with the Circus prototype show that it can improve server throughput by a factor of two or more in workloads with strong file access locality.
Stergios V. Anastasiadis, Rajiv Wickremesinghe, Jeffrey S. Chase
ACM Trans. Storage3
2008 Weighted fair sharing for dynamic virtual clusters
abstract
In a shared server infrastructure, a scheduler controls how quantities of resources are shared over time in a fair manner across multiple, competing consumers. It should support wide (parallel) requests for variable-sized pool of resources, provide assurance of minimum resource allotment on demand, and give predictable assignments. Our approach integrates a fair queuing algorithm with a calendar scheduler. We present WINKS, a proportional share allocation policy that addresses the needs of shared server environments. It extends start-time fair queuing to support wide requests with backfill, advance reservations, dynamic cluster sizing, dynamic request sizing, and intra-flow request prioritization. It also preserves fairness properties across queue transformations and calendar operations needed to implement these extensions.
Laura E. Grit, Jeffrey S. Chase
SIGMETRICS2
2008 Cutting Corners: Workbench Automation for Server Benchmarking
Piyush Shivam, Varun Marupadi, Jeffrey S. Chase, Thileepan Subramaniam, Shivnath Babu
USENIX ATC3
2007 Strong Accountability for Network Storage
Aydan R. Yumerefendi, Jeffrey S. Chase
FAST2
2007 Automated and on-demand provisioning of virtual machines for database applications
abstract
Utility computing delivers compute and storage resources to applications as an 'on-demand utility', much like electricity, from a distributed collection of computing resources. There is great interest in running database applications on utility resources (e.g., Oracle's Grid initiative) due to reduced infrastructure and management costs, higher resource utilization, and the ability to handle sudden load surges. Virtual Machine (VM) technology offers powerful mechanisms to manage a utility resource infrastructure. However, provisioning VMs for applications to meet system performance goals, e.g., to meet service level agreements (SLAs), is an open problem. We are building two systems at Duke - Shirako and NIMO - that collectively address this problem.
Piyush Shivam, Azbayar Demberel, Pradeep Gunda, David Irwin 0001, Laura E. Grit, Aydan R. Yumerefendi, Shivnath Babu, Jeffrey S. Chase
SIGMOD Conference8
2007 Strong accountability for network storage
abstract
This article presents the design, implementation, and evaluation of CATS, a network storage service with strong accountability properties. CATS offers a simple web services interface that allows clients to read and write opaque objects of variable size. This interface is similar to the one offered by existing commercial Internet storage services. CATS extends the functionality of commercial Internet storage services by offering support for strong accountability. A CATS server annotates read and write responses with evidence of correct execution, and offers audit and challenge interfaces that enable clients to verify that the server is faithful. A faulty server cannot conceal its misbehavior, and evidence of misbehavior is independently verifiable by any participant. CATS clients are also accountable for their actions on the service. A client cannot deny its actions, and the server can prove the impact of those actions on the state views it presented to other clients. Experiments with a CATS prototype evaluate the cost of accountability under a range of conditions and expose the primary factors influencing the level of assurance and the performance of a strongly accountable storage server. The results show that strong accountability is practical for network storage systems in settings with strong identity and modest degrees of write-sharing. We discuss how the accountability concepts and techniques used in CATS generalize to other classes of network services.
Aydan R. Yumerefendi, Jeffrey S. Chase
ACM Trans. Storage2
2006 Virtual playgrounds: managing virtual resources in the grid
abstract
Large grid deployments increasingly require abstractions and methods decoupling the work of resource providers and resource consumers to implement scalable management methods. We proposed the abstraction of a virtual workspace (VW) describing a virtual execution environment that can be made dynamically available to authorized grid clients by using well-defined protocols. Virtual workspaces provide resources in controllable ways that are independent of how a resource is consumed. A virtual playground may combine many such workspaces, as well as other aspects of virtual environments, such as networking and storage, to form virtual grids. In this paper, we report on the goals and progress of the virtual playground project and put in context the research to date.
Kate Keahey, Jeffrey S. Chase, Ian T. Foster
IPDPS2
2006 Ensemble-level Power Management for Dense Blade Servers
abstract
One of the key challenges for high-density servers (e.g., blades) is the increased costs in addressing the power and heat density associated with compaction. Prior approaches have mainly focused on reducing the heat generated at the level of an individual server. In contrast, this work proposes power efficiencies at a larger scale by leveraging statistical properties of concurrent resource usage across a collection of systems ("ensemble"). Specifically, we discuss an implementation of this approach at the blade enclosure level to monitor and manage the power across the individual blades in a chassis. Our approach requires low-cost hardware modifications and relatively simple software support. We evaluate our architecture through both prototyping and simulation. For workloads representing 132 servers from nine different enterprise deployments, we show significant power budget reductions at performances comparable to conventional systems.
Parthasarathy Ranganathan, Phil Leech, David Irwin 0001, Jeffrey S. Chase
ISCA4
2006 Grid allocation and reservation - Toward a doctrine of containment: grid hosting with adaptive resource control
abstract
Grid computing environments need secure resource control and predictable service quality in order to be sustainable. We propose a grid hosting model in which independent, self-contained grid deployments run within isolated containers on shared resource provider sites. Sites and hosted grids interact via an underlying resource control plane to manage a dynamic binding of computational resources to containers. We present a prototype grid hosting system, in which a set of independent Globus grids share a network of cluster sites. Each grid instance runs a coordinator that leases and configures cluster resources for its grid on demand. Experiments demonstrate adaptive provisioning of cluster resources and contrast job-level and container-level resource management in the context of two grid application managers.
Lavanya Ramakrishnan, David Irwin 0001, Laura E. Grit, Aydan R. Yumerefendi, Adriana Iamnitchi, Jeffrey S. Chase
SC6
2006 Sharing Networked Resources with Brokered Leases
David Irwin 0001, Jeffrey S. Chase, Laura E. Grit, Aydan R. Yumerefendi, Ken Yocum
USENIX ATC, General Track2
2006 Active and Accelerated Learning of Cost Models for Optimizing Scientific Applications
Piyush Shivam, Shivnath Babu, Jeffrey S. Chase
VLDB3
2005 Lerna: an active storage framework for flexible data access and management
abstract
In the present paper, we examine the problem of supporting application-specific computation within a network file server. Our objectives are (i) to introduce an easy to use yet powerful architecture for executing both custom-developed and legacy applications close to the stored data, (ii) to investigate the performance improvement that we get from data proximity in I/O-intensive processing, and (in) to exploit the I/O-traffic information available within the file server for more effective resource management. One main difference from previous active storage research is our emphasis on the expressive power and usability of the network server interface. We describe an extensible active storage framework that we built in order to demonstrate the feasibility of the proposed system design. We show that accessing large datasets over a wide-area network through a regular file system can penalize the system performance, unless application computation is moved close to the stored data. Our conclusions are substantiated through experimentation with a popular multilayer map warehouse application.
Stergios V. Anastasiadis, Rajiv Wickremesinghe, Jeffrey S. Chase
HPDC3
2005 Making Scheduling "Cool": Temperature-Aware Workload Placement in Data Centers
Justin D. Moore, Jeffrey S. Chase, Parthasarathy Ranganathan, Ratnesh K. Sharma
USENIX ATC, General Track2
2005 Controllable fair queuing for meeting performance goals
Magnus Karlsson 0002, Christos T. Karamanolis, Jeffrey S. Chase
Perform. Evaluation3
2004 Circus: Opportunistic Block Reordering for Scalable Content Servers
Stergios V. Anastasiadis, Rajiv Wickremesinghe, Jeffrey S. Chase
FAST3
2004 Designing for Disasters
Kimberly Keeton, Cipriano A. Santos, Dirk Beyer 0002, Jeffrey S. Chase, John Wilkes
FAST4
2004 Balancing Risk and Reward in a Market-Based Task Service
David Irwin 0001, Laura E. Grit, Jeffrey S. Chase
HPDC3
2004 Globus and PlanetLab Resource Management Solutions Compared
Matei Ripeanu, Mic Bowman, Jeffrey S. Chase, Ian T. Foster, Milan Milenkovic
HPDC3
2004 Correlating Instrumentation Data to System States: A Building Block for Automated Diagnosis and Control
Ira Cohen, Jeffrey S. Chase, Moisés Goldszmidt, Terence Kelly, Julie Symons
OSDI2
2004 Interposed proportional sharing for a storage service utility
abstract
This paper develops and evaluates new share-based scheduling algorithms for differentiated service quality in network services, such as network storage servers. This form of resource control makes it possible to share a server among multiple request flows with probabilistic assurance that each flow receives a specified minimum share of a server's capacity to serve requests. This assurance is important for safe outsourcing of services to shared utilities such as Storage Service Providers.Our approach interposes share-based request dispatching on the network path between the server and its clients. Two new scheduling algorithms are designed to run within an intermediary (e.g., a network switch), where they enforce fair sharing by throttling request flows and reordering requests; these algorithms are adaptations of Start-time Fair Queuing (SFQ) for servers with a configurable degree of internal concurrency. A third algorithm, Request Windows (RW), bounds the outstanding requests for each flow independently; it is amenable to a decentralized implementation, but may restrict concurrency under light load. The analysis and experimental results show that these new algorithms can enforce shares effectively when the shares are not saturated, and that they provide acceptable performance isolation under saturation. Although the evaluation uses a storage service as an example, interposed request scheduling is non-intrusive and views the server as a black box, so it is useful for complex services with no internal support for differentiated service quality.
Jeffrey S. Chase, Jasleen Kaur 0001
SIGMETRICS2
2003 Dynamic Virtual Clusters in a Grid Site Manager
abstract
This paper presents new mechanisms for dynamic resource management in a cluster manager called Cluster-on-Demand (COD). COD allocates servers from a common pool to multiple virtual clusters (vclusters), with independently configured software environments, name spaces, user access controls, and network storage volumes. We present experiments using the popular Sun GridEngine batch scheduler to demonstrate that dynamic virtual clusters are an enabling abstraction for advanced resource management in computing utilities and grids. In particular, they support dynamic, policy-based cluster sharing between local users and hosted Grid services, resource reservation and adaptive provisioning, scavenging of the idle resources, and dynamic instantiation of Grid services. These goals are achieved in a direct and general way through a new set of fundamental cluster management functions, with minimal impact on the Grid middleware itself.
Jeffrey S. Chase, David Irwin 0001, Laura E. Grit, Justin D. Moore, Sara Sprenkle
HPDC1
2003 SHARP: an architecture for secure resource peering
abstract
This paper presents Sharp, a framework for secure distributed resource management in an Internet-scale computing infrastructure. The cornerstone of Sharp is a construct to represent cryptographically protected resource claims ---promises or rights to control resources for designated time intervals---together with secure mechanisms to subdivide and delegate claims across a network of resource managers. These mechanisms enable flexible resource peering : sites may trade their resources with peering partners or contribute them to a federation according to local policies. A separation of claims into tickets and leases allows coordinated resource management across the system while preserving site autonomy and local control over resources. Sharp also introduces mechanisms for controlled, accountable oversubscription of resource claims as a fundamental tool for dependable, efficient resource management. We present experimental results from a Sharp prototype for PlanetLab, and illustrate its use with a decentralized barter economy for global PlanetLab resources. The results demonstrate the power and practicality of the architecture, and the effectiveness of oversubscription for protecting resource availability in the presence of failures.
Yun Fu 0003, Jeffrey S. Chase, Brent N. Chun, Stephen Schwab, Amin Vahdat
SOSP2
2003 Efficient Flow Computation on Massive Grid Terrain Datasets
Lars Arge, Jeffrey S. Chase, Patrick N. Halpin, Laura Toma, Jeffrey Scott Vitter, Dean L. Urban, Rajiv Wickremesinghe
GeoInformatica2
2002 Distributed Computing with Load-Managed Active Storage
abstract
One approach to high-performance processing of massive data sets is to incorporate computation into storage systems. Previous work has shown that this active storage model is effective for a variety of problems. This paper explores opportunities to use active storage as a basis for exploiting asymmetric parallelism in applications using a streaming computation model on collections of fixed-size records. This model is the basis for much of the research in I/O-efficient algorithms, which deals with an important class of massive data problems not studied in previous work on active storage. We present an extension of a streaming computation model for an external memory toolkit to support a flexible mapping of computations to storage-based processors. Our approach enables load-managed active storage: it exposes parallelism, ordering constraints, and primitive computation units to the system, which can configure the application to balance load and make the best use of available processing power Emulation results from a sorting application demonstrate the potential of dynamic adaptation in load-managed active storage.
Rajiv Wickremesinghe, Jeffrey S. Chase, Jeffrey Scott Vitter
HPDC2
2002 Scalability and Accuracy in a Large-Scale Network Emulator
Amin Vahdat, Ken Yocum, Kevin Walsh 0001, Priya Mahadevan, Dejan Kostic, Jeffrey S. Chase
OSDI6
2002 Structure and Performance of the Direct Access File System
Kostas Magoutis, Salimah Addetia, Alexandra Fedorova, Margo I. Seltzer, Jeffrey S. Chase, Andrew J. Gallatin, Richard Kisley, Rajiv Wickremesinghe, Eran Gabber
USENIX ATC, General Track5
2002 The Trickle-Down Effect: Web Caching and Server Request Distribution
Ronald P. Doyle, Jeffrey S. Chase, Syam Gadde, Amin Vahdat
Comput. Commun.2
2002 Interposed request routing for scalable network storage
abstract
This paper explores interposed request routing in Slice, a new storage system architecture for high-speed networks incorporating network-attached block storage. Slice interposes a request switching filter---called a μproxy---along each client's network path to the storage service (e.g., in a network adapter or switch). The μproxy intercepts request traffic and distributes it across a server ensemble. We propose request routing schemes for I/O and file service traffic, and explore their effect on service structure. The Slice prototype uses a packet filter μproxy to virtualize the standard Network File System (NFS) protocol, presenting to NFS clients a unified shared file volume with scalable bandwidth and capacity. Experimental results from the industry-standard SPECsfs97 workload demonstrate that the architecture enables construction of powerful network-attached storage services by aggregating cost-effective components on a switched Gigabit Ethernet LAN.
Darrell C. Anderson, Jeffrey S. Chase, Amin Vahdat
ACM Trans. Comput. Syst.2
2001 Energy Management for Server Clusters
abstract
The central point of this paper is that energy should be viewed as an important element of resource management for Web sites, hosting centers, and other Internet server clusters. In particular, we are developing a system to manage server resources so that cluster power demand scales with request throughput. This can yield significant energy savings because server clusters are sized for peak load, while traces show that traffic varies by factors of 3-6 or more through any day or week, with average load often less than 50% of peak. We propose energy-conscious service provisioning, in which the system continuously monitors load and adaptively provisions server capacity. This promises both economic and environmental benefits.
Jeffrey S. Chase, Ronald P. Doyle
HotOS1
2001 Anypoint Communication Protocol
abstract
Summary form only given. We are developing the Anypoint Communication Protocol (ACP). ACP clients establish connections to abstract services, represented at the network edge by Anypoint intermediaries. The intermediary is an intelligent network switch that acts as an extension of the service; it encapsulates a service-specific policy for distributing requests among servers in the active set for each service. The switch routes incoming requests on each ACP connection to any active server at the discretion of the service routing policy, hence the name "Anypoint". The Anypoint abstraction and ACP protocol enable virtualization using intermediaries for a general class of wide-area network services based on request/response communication over persistent transport connections. Potential applications include scalable IP-based network storage protocols and next-generation Web services.
Ken Yocum, Jeffrey S. Chase, Amin Vahdat
HotOS2
2001 Managing Energy and Server Resources in Hosting Centres
abstract
Internet hosting centers serve multiple service sites from a common hardware base. This paper presents the design and implementation of an architecture for resource management in a hosting center operating system, with an emphasis on energy as a driving resource management issue for large server clusters. The goals are to provision server resources for co-hosted services in a way that automatically adapts to offered load, improve the energy efficiency of server clusters by dynamically resizing the active server set, and respond to power supply disruptions or thermal events by degrading service in accordance with negotiated Service Level Agreements (SLAs).Our system is based on an economic approach to managing shared server resources, in which services "bid" for resources as a function of delivered performance. The system continuously monitors load and plans resource allotments by estimating the value of their effects on service performance. A greedy resource allocation algorithm adjusts resource prices to balance supply and demand, allocating resources to their most efficient use. A reconfigurable server switching infrastructure directs request traffic to the servers assigned to each service. Experimental results from a prototype confirm that the system adapts to offered load and resource availability, and can reduce server energy usage by 29% or more for a typical Web workload.
Jeffrey S. Chase, Darrell C. Anderson, Prachi N. Thakar, Amin Vahdat, Ronald P. Doyle
SOSP1
2001 Payload Caching: High-Speed Data Forwarding for Network Intermediaries
Ken Yocum, Jeffrey S. Chase
USENIX ATC, General Track2
2001 Web caching and content distribution: a view from the interior
Syam Gadde, Jeffrey S. Chase, Michael Rabinovich
Comput. Commun.2
2000 Failure-Atomic File Access in an Interposed Network Storage System
abstract
Presents a recovery protocol for block I/O operations in Slice, a storage system architecture for high-speed LANs incorporating network-attached block storage. The goal of the Slice architecture is to provide a network file service with scalable bandwidth and capacity while preserving compatibility with off-the-shelf clients and file server appliances. The Slice prototype "virtualizes" the Network File System (NFS) protocol by interposing a request switching filter at the client's interface to the network storage system (e.g. in a network adapter or switch). The distributed Slice architecture separates functions that are typically combined in central file servers, introducing new challenges for failure atomicity. This paper presents a protocol for atomic file operations and recovery in the Slice architecture, and related support for reliable file storage using mirrored striping. Experimental results from the Slice prototype show that the protocol has low cost in the common case, allowing the system to deliver client file access bandwidths approaching Gbit/s network speeds.
Darrell C. Anderson, Jeffrey S. Chase
HPDC2
2000 Interposed Request Routing for Scalable Network Storage
Darrell C. Anderson, Jeffrey S. Chase, Amin Vahdat
OSDI2
1999 Potentials and Limitations of Fault-Based Markov Prefetching for Virtual Memory Pages
abstract
No abstract available.
Gretta Bartels, Anna R. Karlin, Darrell C. Anderson, Jeffrey S. Chase, Henry M. Levy, Geoffrey M. Voelker
SIGMETRICS4
1998 Implementing Cooperative Prefetching and Caching in a Globally-Managed Memory System
abstract
This paper presents cooperative prefetching and caching --- the use of network-wide global resources (memories, CPUs, and disks) to support prefetching and caching in the presence of hints of future demands. Cooperative prefetching and caching effectively unites disk-latency reduction techniques from three lines of research: prefetching algorithms, cluster-wide memory management, and parallel I/O. When used together, these techniques greatly increase the power of prefetching relative to a conventional (non-global-memory) system. We have designed and implemented PGMS, a cooperative prefetching and caching system, under the Digital Unix operating system running on a 1.28 Gb/sec Myrinet-connected cluster of DEC Alpha workstations. Our measurements and analysis show that by using available global resources, cooperative prefetching can obtain significant speedups for I/O-bound programs. For example, for a graphics rendering application, our system achieves a speedup of 4.9 over a non-prefetching version of the same program, and a 3.1-fold improvement over that program using local-disk prefetching alone.
Geoffrey M. Voelker, Eric J. Anderson, Tracy Kimbrel, Michael J. Feeley, Jeffrey S. Chase, Anna R. Karlin, Henry M. Levy
SIGMETRICS5
1998 Cheating the I/O Bottleneck: Network Storage with Trapeze/Myrinet
Darrell C. Anderson, Jeffrey S. Chase, Syam Gadde, Andrew J. Gallatin, Ken Yocum, Michael J. Feeley
USENIX ATC2
1998 Not all Hits are Created Equal: Cooperative Proxy Caching Over a Wide-Area Network
Michael Rabinovich, Jeffrey S. Chase, Syam Gadde
Comput. Networks2
1997 Cut-Through Delivery in Trapeze: An Exercise in Low-Latency Messaging
abstract
New network technology continues to improve both the latency and bandwidth of communication in computer clusters. The fastest high-speed networks approach or exceed the I/O bus bandwidths of "gigabit-ready" hosts. These advances introduce new considerations for the design of network interfaces and messaging systems for low-latency communication. This paper investigates cut-through delivery, a technique for overlapping host I/O DMA transfers with network traversal. Cut-through delivery significantly reduces end-to-end latency of large messages, which are often critical for application performance. We have implemented cut-through delivery in Trapeze, a new messaging substrate for network memory and other distributed operating system services. Our current Trapeze prototype is capable of demand-fetching 8 K virtual memory pages in 200 /spl mu/s across a Myrinet cluster of DEC AlphaStations.
Ken Yocum, Jeffrey S. Chase, Andrew J. Gallatin, Alvin R. Lebeck
HPDC2
1994 Integrating Coherency and Recoverability in Distributed Systems
Michael J. Feeley, Jeffrey S. Chase, Vivek R. Narasayya, Henry M. Levy
OSDI2
1994 Sharing and Protection in a Single-Address-Space Operating System
abstract
This article explores memory sharing and protection support in Opal, a single-address-space operating system designed for wide-address (64-bit) architectures. Opal threads execute within protection domains in a single shared virtual address space. Sharing is simplified, because addresses are context independent. There is no loss of protection, because addressability and access are independent; the right to access a segment is determined by the protection domain in which a thread executes. This model enables beneficial code-and data-sharing patterns that are currently prohibitive, due in part to the inherent restrictions of multiple address spaces, and in part to Unix programming style. We have designed and implemented an Opal prototype using the Mach 3.0 microkernel as a base. Our implementation demonstrates how a single-address-space structure can be supported alongside of other environments on a modern microkernel operating system, using modern wide-address architectures. This article justifies the Opal model and its goals for sharing and protection, presents the system and its abstractions, describes the prototype implementation, and reports experience with integrated applications.
Jeffrey S. Chase, Henry M. Levy, Michael J. Feeley, Edward D. Lazowska
ACM Trans. Comput. Syst.1
1992 Architectural Support for Single Address Space Operating Systems
abstract
Recent microprocessor announcements show a trend toward wide-address computers: architectures that support 64 bits of virtual address space.Such architectures facilitate fundamentally new operating system organizations that promote efficient data sharing and cooperation, both between complex applications and between parts of the operating system itself.Machinery.
Eric J. Koldinger, Jeffrey S. Chase, Susan J. Eggers
ASPLOS2
1992 Lightweight Shared Objects in a 64-Bit Operating System
abstract
Object-oriented models are a popular basis for supporting uniform sharing of data and services in operating systems, distributed programming systems, and database systems. We term systems that use objects for these purposes object sharing systems. Operating systems in common use have nonuniform addressing models, making the uniform object naming required by object sharing systems expensive and difficult to implement. We argue that emerging 64-bit architectures make it practical to support uniform naming at the virtual addressing level, eliminating a key implementation problem for object sharing systems. We describe facilities for object-based sharing of persistent data and services in Opal, an operating system we are developing for paged 64-bit architectures. The distinctive feature of Opal is that object This paper will appear in identical form in the proceedings of the Conference on Object-Oriented Programming Systems, Languages, and Applications (OOPSLA), October 1992. This work w...
Jeffrey S. Chase, Henry M. Levy, Edward D. Lazowska, Miche Baker-Harvey
OOPSLA1
1991 Dynamic Node Reconfiguration in a Parallel-Distributed Environment
abstract
Idle workstationsin a network represent a significant computing potential.In particular, their processing power can be used by parallel-distributed programs that treat the network as a loosely-coupled multiprocessor.Our experiments with Amber show that node reconfiguration can be implemented easily and efficiently in a runtime library.work speed will substantially reduce the cost
Michael J. Feeley, Brian N. Bershad, Jeffrey S. Chase, Henry M. Levy
PPoPP3
1989 The Amber System: Parallel Programming on a Network of Multiprocessors
abstract
This paper describes a programming system called Amber that permits a single application program to use a homogeneous network of computers in a uniform way, making the network appear to the application as an integrated multiprocessor. Amber is specifically designed for high performance in the case where each node in the network is a shared-memory multiprocessor.
Jeffrey S. Chase, Franz G. Amador, Edward D. Lazowska, Henry M. Levy, Richard J. Littlefield
SOSP1