EDBT 2026 Demo / reviewers in the wild / expert
Chandra Krintz
dblp:87/1311
· DBLP profile ↗
70ranked-venue papers
11as first author
11since 2021 · last 2024
0000-0003-4972-0669ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 21 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distributed Dataflow Across the Edge-Cloud ContinuumabstractInternet of Things (IoT) applications span the edge-cloud continuum to form multiscale distributed systems. The heterogeneity that defines this architecture, coupled with the asynchronous, event-triggered and failure-prone nature of these deployments create significant programming and maintenance challenges for developers of IoT applications. To address this impediment to innovation, we present Lam-1nar’a dataflow programming model for IoT applications implemented using a novel log-based and concurrent runtime system that spans all resource scales. We describe the properties that underpin Laminar'sdesign and compare it to a lower-level event-based approach. We show that Laminar'sdataflow model hides many of the complexities of “lock-free” event-driven programming. Through an empirical evaluation of Laminar,we find its design and implementation are both more straightforward for developers and more performant. Tyler Ekaireb, Lukas Brand, Nagarjun Avaraddy, Markus Mock, Chandra Krintz, Richard Wolski |
CLOUD | 5 |
| 2024 | Energy-Aware IoT Deployment PlanningabstractIncreasingly, the Internet of Things (IoT) is evolving toward an architecture consisting of sensing and actuation devices communicating with edge computers and storage systems. These "edge deployments" localize communication, computation, and storage for security, increased efficiencies (e.g. lower latency response), and reliability. In settings where electrical power infrastructure is lacking, however, these deployments typically rely on renewable energy and battery storage for power. Peiyuan Guan, Animesh Dangwal, Amirhosein Taherkordi, Richard Wolski, Chandra Krintz |
CF | 5 |
| 2023 | Depot: Dependency-Eager Platform of TransformationsabstractThis paper presents a new model for a data management system specifically designed to enable community-curated data repositories and collaboration. Depot (a Dependence-Eager Platform of Transformations) is based on a data-lake approach that eases the technological burdens associated with data contribution while providing an interactive programming environment for developing transformations that result in structured tables supporting SQL database operations. Crucially, Depot implements lazy evaluation of these transformations so that only the structured data that is demanded by a data consumer is generated. Until the structured data is "materialized," Depot tracks and maintains the dependencies that are required to perform the eventual materialization. This lazy approach to creating structured data allows Depot to maintain a smaller resource footprint compared to a typical data warehouse approach while maintaining the flexibility of the data lake model. Furthermore, Depot is designed as a community-sustainable platform. The initial prototype is implemented for cloud deployment and it distributes the storage and ETL workload cost among the data consumers. Performance results of the early prototype are encouraging, making Depot a new infrastructure for creating data lakes that foster contributed-consumer collaboration. Kerem Çelik, Samridhi Maheshwari, Shereen ElSayed, Markus Mock, Chandra Krintz, Richard Wolski |
CloudCom | 5 |
| 2023 | GreenCoin: A Renewable Energy-Aware CryptocurrencyabstractIn this paper, we propose GreenCoin – an energy-efficient cryptocurrency system with mining protocols designed to favor locations with relatively higher availability of renewable energy. Traditionally, crypto coin mining involves solving complex mathematical problems by high-end computing devices consuming an enormous amount of electricity, thus adversely affecting net carbon emissions. To reduce cost and emissions, GreenCoin uses a modified proof of stake (PoS) consensus algorithm, which itself is more energy efficient compared to other state-of-the-art methods. Our modified PoS algorithm, called Green PoS (GPoS), allows GreenCoin to favor nodes (with reward and privilege) located in regions with higher availability of renewable energy. We present a detailed system architecture of GreenCoin and explain the operating method of GPoS. We also provide results from empirical studies demonstrating the renewable energy-aware approach of GreenCoin. Shivaansh Kapoor, Chandra Krintz, Richard Wolski, Markus Mock |
IC2E | 3 |
| 2023 | Data Acquisition and Analysis for Improving the Utility of Low Cost Soil Moisture SensorsabstractTo cultivate healthy plants and high crop yields, growers must be able to measure soil moisture and irrigate accordingly. Errors in soil moisture measurements can lead to irrigation mismanagement with costly consequences. In this paper, we present a new approach to smart computing for irrigation management to address these challenges at a lower cost. We calibrate low cost, low precision soil moisture sensors to more accurately distinguish wet from dry soils using high cost, high precision Davis Instrument sensors. We investigate different modeling techniques including the natural log of the odds ratio (Log-odds), Monte Carlo simulation, and linear regression to distinguish between wet and moist soils and to establish a trustworthy threshold between these two moisture states. We have also developed a new smartphone application that simplifies the process of data collection and implements our analysis approach. The application is extensible by others and provides growers with low cost, data-driven decision support for irrigation. We implement our approach for UCSB’s Edible Campus student farm and empirically evaluate it using multiple test beds. Our results show an accuracy rate of 91% and lowers costs by 4x per deployment, making it useful for gardeners and farmers alike. Gautam Mundewadi, Richard Wolski, Chandra Krintz |
SMARTCOMP | 3 |
| 2023 | Replicated Versioned Data Structures for Wide-Area Distributed SystemsabstractIn this work, we investigate the integration of replicated versioned data structures and append-only distributed storage systems. Doing so facilitates high availability and scalability while providing developer access to different versions of program data structures across program executions. Modern distributed systems such as the Internet of Things (IoT) often employ multi-tiered (cloud/edge/sensors) architectures consisting of a wide array of heterogeneous devices generating data frequently. Hence system availability is imperative to avoid data loss, while scalability is required for the efficient operation of the system not only within the same tier but across different tiers as well. Our proposed approach replicates, persists, and versionsprogram data structuressuch as binary search trees and linked lists for use in distributed IoT applications. The versioning and persistence of these structures aid failure recovery and facilitate system debugging from its inception instead of making such considerations an afterthought. Moreover, our experiments suggest versioned data structures can perform better in applications performing high volumes of temporal queries versus traditional methods of persisting data (e.g., in a database). We empirically evaluate the overheads associated with versioning and storage persistence of program data structures, present experimental results for multiple end-to-end applications, and demonstrate the scalability of this approach. Chandra Krintz, Richard Wolski |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Log-Based CRDT for Edge ApplicationsabstractIn this paper, we investigate extensions for Conflict-Free Replicated Data Types (CRDTs) that permit their use in failure-prone, heterogeneous, resource-constrained, distributed, multi-tier (cloud/edge/device) cloud deployments such as the Internet-of-Things (IoT), while addressing multiple CRDT limitations. Specifically, we employ distributed logging to implement robust, strong eventual consistency of replicas. Our approach also enables uniform reversal of operations and precludes the requirement of exactly-once delivery and idempotence imposed by operation-based CRDTs. Moreover, it exposes CRDT versions for use in debugging and history-based programming. We evaluate our approach for commonly used CRDTs and show that it enables higher operation throughput (up to 1.8x) versus conventional CRDTs for the workloads we consider. Chandra Krintz, Richard Wolski |
IC2E | 2 |
| 2021 | On the Future of Cloud EngineeringabstractEver since the commercial offerings of the Cloud started appearing in 2006, the landscape of cloud computing has been undergoing remarkable changes with the emergence of many different types of service offerings, developer productivity enhancement tools, and new application classes as well as the manifestation of cloud functionality closer to the user at the edge. The notion of utility computing, however, has remained constant throughout its evolution, which means that cloud users always seek to save costs of leasing cloud resources while maximizing their use. On the other hand, cloud providers try to maximize their profits while assuring service-level objectives of the cloud-hosted applications and keeping operational costs low. All these outcomes require systematic and sound cloud engineering principles. The aim of this paper is to highlight the importance of cloud engineering, survey the landscape of best practices in cloud engineering and its evolution, discuss many of the existing cloud engineering advances, and identify both the inherent technical challenges and research opportunities for the future of cloud computing in general and cloud engineering in particular. David Bermbach, Abhishek Chandra, Chandra Krintz, Aniruddha S. Gokhale, Aleksander Slominski, Lauritz Thamsen, Everton Cavalcante, Tian Guo 0001, Ivona Brandic, Richard Wolski |
IC2E | 3 |
| 2021 | PEDaLS: Persisting Versioned Data StructuresabstractIn this paper, we investigate how to automatically persist versioned data structures in distributed settings (e.g. cloud + edge) using append-only storage. By doing so, we facilitate resiliency by enabling program state to survive program activations and termination, and program-level data structures and their version information to be accessed programmatically by multiple clients (for replay, provenance tracking, debugging, and coordination avoidance, and more). These features are useful in distributed, failure-prone contexts such as those for heterogeneous and pervasive Internet of Things (IoT) deployments. We prototype our approach within an open-source, distributed operating system for IoT. Our results show that it is possible to achieve algorithmic complexities similar to those of in-memory versioning but in a distributed setting. Chandra Krintz, Richard Wolski |
IC2E | 2 |
| 2021 | CAPLets: Resource Aware, Capability-Based Access Control for IoT
Fatih Bakir, Chandra Krintz, Richard Wolski |
SEC | 2 |
| 2021 | Edge-adaptable serverless acceleration for machine learning Internet of Things applicationsabstractAbstract Serverless computing is an emerging event‐driven programming model that accelerates the development and deployment of scalable web services on cloud computing systems. Though widely integrated with the public cloud, serverless computing use is nascent for edge‐based, Internet of Things (IoT) deployments. In this work, we present STOIC (serverless teleoperable hybrid cloud), an IoT application deployment and offloading system that extends the serverless model in three ways. First, STOIC adopts a dynamic feedback control mechanism to precisely predict latency and dispatch workloads uniformly across edge and cloud systems using a distributed serverless framework. Second, STOIC leverages hardware acceleration (e.g., GPU resources) for serverless function execution when available from the underlying cloud system. Third, STOIC can be configured in multiple ways to overcome deployment variability associated with public cloud use. We overview the design and implementation of STOIC and empirically evaluate it using real‐world machine learning applications and multitier IoT deployments (edge and cloud). Specifically, we show that STOIC can be used fortrainingimage processing workloads (for object recognition)—once thought too resource‐intensive for edge deployments. We find that STOIC reduces overall execution time (response latency) and achieves placement accuracy that ranges from 92% to 97%. Chandra Krintz, Richard Wolski |
Softw. Pract. Exp. | 2 |
| 2020 | NanoLambda: Implementing Functions as a Service at All Resource Scales for the Internet of ThingsabstractInternet of Things (IoT) devices are becoming increasingly prevalent in our environment, yet the process of programming these devices and processing the data they produce remains difficult. Typically, data is processed on device, involving arduous work in low level languages, or data is moved to the cloud, where abundant resources are available for Functions as a Service (FaaS) or other handlers. FaaS is an emerging category of flexible computing services, where developers deploy self-contained functions to be run in portable and secure containerized environments; however, at the moment, these functions are limited to running in the cloud or in some cases at the “edge” of the network using resource rich, Linux-based systems.In this paper, we present NanoLambda, a portable platform that brings FaaS, high-level language programming, and familiar cloud service APIs to non-Linux and microcontroller-based IoT devices. To enable this, NanoLambda couples a new, minimal Python runtime system that we have designed for the least capable end of the IoT device spectrum, with API compatibility for AWS Lambda and S3. NanoLambda transfers functions between IoT devices (sensors, edge, cloud), providing power and latency savings while retaining the programmer productivity benefits of high-level languages and FaaS. A key feature of NanoLambda is a scheduler that intelligently places function executions across multi-scale IoT deployments according to resource availability and power constraints. We evaluate a range of applications that use NanoLambda to run on devices as small as the ESP8266 with 64KB of ram and 512KB flash storage. Gareth George, Fatih Bakir, Richard Wolski, Chandra Krintz |
SEC | 4 |
| 2020 | Detecting Performance Anomalies in Cloud Platform ApplicationsabstractWe present Roots, a full-stack monitoring and analysis system for performance anomaly detection and bottleneck identification in cloud platform-as-a-service (PaaS) systems. Roots facilitates application performance monitoring as a core capability of PaaS clouds, and relieves the developers from having to instrument application code. Roots tracks HTTP/S requests to hosted cloud applications and their use of PaaS services. To do so it employs lightweight monitoring of PaaS service interfaces. Roots processes this data in the background using multiple statistical techniques that in combination detect performance anomalies (i.e. violations of service-level objectives). For each anomaly, Roots determines whether the event was caused by a change in the request workload or by a performance bottleneck in a PaaS service. By correlating data collected across different layers of the PaaS, Roots is able to trace high-level performance anomalies to bottlenecks in specific components in the cloud platform. We implement Roots using the AppScale PaaS and evaluate its overhead and accuracy. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | Seneca: Fast and Low Cost Hyperparameter Search for Machine Learning ModelsabstractThe goal of our work is to simplify and expedite the construction and evaluation of machine learning models using autoscaled cloud computing resources. To enable this, we develop an open source system called Seneca, which leverages the serverless programming model and its implementation in Amazon Web Services (AWS) Lambda. Seneca takes a machine learning application, dataset, and a list of possible hyperparameter options as input and automatically constructs an AWS Lambda function. The function ingresses and splits the input dataset into training and testing subsets and constructs, tests, and evaluates (i.e. scores) a machine learning model for a given set of hyperparameter values. Seneca concurrently invokes functions for all combinations of the hyperparameters specified. It then returns the configuration (or model) that results in the best score to the user. In this paper, we overview the design and implementation of Seneca, and empirically evaluate its performance for a popular classification application. Chandra Krintz, Markus Mock, Richard Wolski |
CLOUD | 2 |
| 2019 | Analyzing AWS Spot Instance PricingabstractMany cloud computing vendors offer a preemptible class of service for rented virtual machines. In November 2017, Amazon.com changed the pricing mechanism for its preemptible "spot instances" so that prices would change more "smoothly." This paper analyzes the effect of this change on spot instance prices. It examines the prices immediately before and after the mechanism change to determine the extent to which prices themselves changed. It then compares the 90-day period immediately after the change in mechanism to the next 90-day period. Finally, it compares the two most recent 90-day periods (ending on October 15, 2018). Our results indicate that in addition to smoothing prices, the mechanism change introduced generally higher prices which is a trend that continues. Gareth George, Richard Wolski, Chandra Krintz, John Brevik |
IC2E | 3 |
| 2018 | Tracing Function Dependencies across CloudsabstractIn this paper, we present Lowgo, a crosscloud tracing tool for capturing causal relationships in serverless applications. To do so, Lowgo records dependencies between functions, through cloud services, and across regions to facilitate debugging and reasoning about highly concurrent, multi-cloud applications. We empirically evaluate Lowgo using microbenchmarks and multi-function and multi-cloud applications. We find that Lowgo is able to capture causal dependencies with overhead that ranges from 2-12%, which is less than half that of the best-performing, cloud-specific approach. Wei-Tsung Lin, Chandra Krintz, Richard Wolski |
IEEE CLOUD | 2 |
| 2018 | Tracking Causal Order in AWS Lambda ApplicationsabstractServerless computing is a new cloud programming and deployment paradigm that is receiving wide-spread uptake. Serverless offerings such as Amazon Web Services (AWS) Lambda, Google Functions, and Azure Functions automatically execute simple functions uploaded by developers, in response to cloud-based event triggers. The serverless abstraction greatly simplifies integration of concurrency and parallelism into cloud applications, and enables deployment of scalable distributed systems and services at very low cost. Although a significant first step, the serverless abstraction requires tools that software engineers can use to reason about, debug, and optimize their increasingly complex, asynchronous applications. Toward this end, we investigate the design and implementation of GammaRay, a cloud service that extracts causal dependencies across functions and through cloud services, without programmer intervention. We implement GammaRay for AWS Lambda and evaluate the overheads that it introduces for serverless micro-benchmarks and applications written in Python. Wei-Tsung Lin, Chandra Krintz, Richard Wolski, Xiaogang Cai, Tongjun Li, Weijin Xu |
IC2E | 2 |
| 2017 | PYTHIA: Admission Control for Multi-Framework, Deadline- Driven, Big Data WorkloadsabstractIn this paper, we present PYTHIA, deadline-aware admission control for systems that execute jobs from multiple big data (batch) frameworks using shared resources. PYTHIA adds support for deadline-driven workloads in resource-constrained cloud settings, for use by resource negotiators such as Apache Mesos or YARN. PYTHIA uses histories of job statistics to estimate the minimum number of CPUs to allocate to a job in order for it to meet its deadline. PYTHIA admits jobs when these resources are available. Any job not admitted “fails fast”and wastes no resources. We implement a PYTHIA prototype and empirically evaluate it using production YARN traces under different resource constraints and deadline assignments. Our results show that PYTHIA is able to meet significantly more deadlines than fair share approaches and wastes fewer cloud resources in resource-limited scenarios, for the workloads, cluster sizes, and deadline assignments that we consider Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
CLOUD | 2 |
| 2017 | Justice: A Deadline-Aware, Fair-Share Resource Allocator for Implementing Multi-AnalyticsabstractIn this paper, we present Justice, a fair-share deadline-aware resource allocator for big data cluster managers. In resource constrained environments, where resource contention introduces significant execution delays, Justice outperforms the popular existing fair-share allocator that is implemented as part of Mesos and YARN. Justice uses deadline information supplied with each job and historical job execution logs to implement admission control. It automatically adapts to changing workload conditions to assign enough resources for each job to meet its deadline "just in time." We use trace-based simulation of production YARN workloads to evaluate Justice under different deadline formulations. We compare Justice to the existing fair-share allocation policy deployed on cluster managers like YARN and Mesos and find that in resource-constrained settings, Justice improves fairness, satisfies significantly more deadlines, and utilizes resources more efficiently. Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
CLUSTER | 2 |
| 2017 | EXFed: Efficient Cross-Federation with Availability SLAs on Preemptible IaaS InstancesabstractPrivate IaaS clouds offer the benefits of cloud computing on-site but their efficiency is limited by capacity constraints during peak times. We present EXFed, an efficient cross-federation system for IaaS clouds that "ships" jobs between clouds and provides ahead-of-time certainty about resource availability despite retaining individual clouds' ability to preempt foreign workload after admission. Clouds participating in the federation remain in control of their local resources at all times and exclusively use a predictable tier of preemptible instances to run federated jobs. This predictable tier is enabled through a new method that provides an SLA on the preemption probability of groups of instances. The SLA is learned statistically from cloud utilization and data transfer rates in the recent past. We deploy EXFed across multiple data centers and evaluate its robustness under realistic and adverse scenarios with production traces recorded from industrial "big data" clouds. Alexander Pucher, Richard Wolski, Chandra Krintz |
IC2E | 3 |
| 2017 | Performance Monitoring and Root Cause Analysis for Cloud-hosted Web ApplicationsabstractIn this paper, we describe Roots - a system for automatically identifying the "root cause" of performance anomalies in web applications deployed in Platform-as-a-Service (PaaS) clouds. Roots does not require application-level instrumentation. Instead, it tracks events within the PaaS cloud that are triggered by application requests using a combination of metadata injection and platform-level instrumentation. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
WWW | 2 |
| 2016 | Big data framework interference in restricted private cloud settingsabstractIn this paper, we characterize the behavior of “big” and “fast” data analysis frameworks, in multi-tenant, shared settings for which computing resources (CPU and memory) are limited, an increasingly common scenario used to increase utilization and lower cost. We study how popular analytics frameworks behave and interfere with each other under such constraints. We empirically evaluate Hadoop, Spark, and Storm multi-tenant workloads managed by Mesos. Our results show that in constrained environments, there is significant performance interference that manifests in failed fair sharing, performance variability, and deadlock of resources. Stratos Dimopoulos, Chandra Krintz, Richard Wolski |
IEEE BigData | 2 |
| 2016 | Stochastic Simulation Service: Bridging the Gap between the Computational Expert and the BiologistabstractWe present StochSS: Stochastic Simulation as a Service, an integrated development environment for modeling and simulation of both deterministic and discrete stochastic biochemical systems in up to three dimensions. An easy to use graphical user interface enables researchers to quickly develop and simulate a biological model on a desktop or laptop, which can then be expanded to incorporate increasing levels of complexity. StochSS features state-of-the-art simulation engines. As the demand for computational power increases, StochSS can seamlessly scale computing resources in the cloud. In addition, StochSS can be deployed as a multi-user software environment where collaborators share computational resources and exchange models via a public model repository. We demonstrate the capabilities and ease of use of StochSS with an example of model development and simulation at increasing levels of complexity. Brian Drawert, Andreas Hellander, Benjamin B. Bales, Debjani Banerjee, Giovanni Bellesia, Bernie J. Daigle Jr., Geoffrey Douglas, Mengyuan Gu, Anand Gupta, Stefan Hellander, Christopher B. Horuk, Dibyendu Nath, Aviral Takkar, Sheng Wu 0002, Per Lötstedt, Chandra Krintz, Linda R. Petzold |
PLoS Comput. Biol. | 16 |
| 2015 | Response time service level agreements for cloud-hosted web applicationsabstractCloud computing is a successful model for hosting web-facing applications that are accessed by their users as services. While clouds currently offer Service Level Agreements (SLAs) containing guarantees of availability, they do not make performance guarantees for deployed applications. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
SoCC | 2 |
| 2015 | Service-Level Agreement Durability for Web Service Response TimeabstractCloud computing is an attractive model for deploying web services in a highly scalable manner. Users access such cloud-hosted services via their web-facing application programming interfaces (APIs). Prior work has shown that it is possible to use a combined approach of static analysis and cloud platform monitoring to predict the response time upper bounds of such web APIs. This technique can be employed to automatically generate service level agreements (SLAs) concerning the performance of cloud-hosted web APIs. In this work, we explore the validity period of auto-generated SLAs in cloud settings. We discuss a simple model by which API consumers can establish a response time SLA with the cloud platform, and renegotiate it when/if the SLA becomes invalid due to the dynamic nature of the cloud. Using empirical methods and simulations on a real world public cloud platform, we show that it is possible to auto-generatehighly durable response time SLAs for cloud-hosted web APIs, thereby keeping the number of SLA invalidations and renegotiations very low, over long periods. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
CloudCom | 2 |
| 2015 | SuperContra: Cross-Language, Cross-Runtime Contracts as a ServiceabstractThis paper presents SuperContra - a Design-by-Contract (DbC) framework that can ship with future PaaS offerings to enforce lightweight contracts across different programming systems, as-a-service. SuperContra is unique in that developers employ a familiar, high-level language to write contracts regardless of the programming language used to implement the component under test. We evaluate SuperContra using widely used, open-source software and compare its performance against existing DbC frameworks. Our results show that SuperContra performs on par with non-service-based DbC approaches and in some cases similarly to code running without contracts. Stratos Dimopoulos, Chandra Krintz, Richard Wolski, Anand Gupta |
IC2E | 2 |
| 2015 | EAGER: Deployment-Time API Governance for Modern PaaS CloudsabstractTo track, control, and compel reuse of web APIs, we investigate a new approach to API governance -- combined policy, implementation, and deployment control of web APIs. Our approach, called EAGER, provides a software architecture that integrates into PaaS platforms to support systemwide, deployment-time enforcement of governance policies. Specifically, EAGER checks for and prevents backward incompatible API changes from being deployed into production PaaS clouds, enforces service reuse, and facilitates enforcement of other best practices in software maintenance via policies. Our experiments with an EAGER prototype show that enforcing API governance at deployment-time in PaaS clouds is efficient and scalable to thousands of APIs and policies. Hiranya Jayathilaka, Chandra Krintz, Richard Wolski |
IC2E | 2 |
| 2015 | Using Trustworthy Simulation to Engineer Cloud SchedulersabstractIn recent years, researchers have contributed promising new techniques for allocating cloud resources in more robust, efficient, and ecologically sustainable ways. Unfortunately, the wide-spread use of these techniques in production systems has, to date, remained elusive. One reason for this is that the state of the art for investigating these innovations at scale often relies solely on model-driven simulation. Production-grade cloud software, however, demands certainty and precision for development and business planning that only comes from validating simulation against empirical observation. In this work, we take an alternative approach to facilitating cloud research and engineering in order to transition innovations to production deployment faster. In particular, we present a new methodology that complements existing model-driven simulation with platform-specific and statistically trustworthy results. We simulate systems at scales and on time frames that are testable, and then, based on the statistical validation of these simulations, investigate scenarios beyond those feasibly observable in practice. We demonstrate the approach by developing an energy-aware cloud scheduler and evaluating it using production and synthetic traces in faster than real time. Our results show that we can accurately simulate a production IaaS system, ease capacity planning, and expedite the reliable development of its components and extensions. Alexander Pucher, Emre Gul, Richard Wolski, Chandra Krintz |
IC2E | 4 |
| 2014 | Cloud Platform Support for API GovernanceabstractAs scalable information technology evolves to a more cloud-like model, digital assets (code, data and software environments) increasingly require curation as web-accessible services. "Service-izing" digital assets consists of encapsulating assets in software that exposes them to web and mobile applications via well-defined yet flexible, network accessible, application programming interfaces (APIs). In this paper, we postulate that recent advances in cloud computing make cloud platforms as-a-service (PaaS) ideal for deployment, lifecycle management, and policy-based control i.e. API governance - for extant and future digital assets. Toward this end, we overview API governance as a PaaS technology and outline some early results generated by our investigation of a prototype we are developing, called EAGER, for implementing API governance at scale. Chandra Krintz, Hiranya Jayathilaka, Stratos Dimopoulos, Alexander Pucher, Richard Wolski, Tevfik Bultan |
IC2E | 1 |
| 2013 | Cloud Platform Datastore Support
Navraj Chohan, Chris Bunch, Chandra Krintz, Navyasri Canumalla |
J. Grid Comput. | 3 |
| 2012 | Language and Runtime Support for Automatic Configuration and Deployment of Scientific Computing Software over Cloud Fabrics
Chris Bunch, Brian Drawert, Navraj Chohan, Chandra Krintz, Linda R. Petzold, Khawaja S. Shams |
J. Grid Comput. | 4 |
| 2011 | Database-Agnostic Transaction Support for Cloud InfrastructuresabstractIn this paper, we present and empirically evaluate the performance of database-agnostic transaction (DAT) support for the cloud. Our design and implementation of DAT is scalable, fault-tolerant, and requires only that the data store provide atomic, row-level access. Our approach enables applications to employ a single transactional data store API that can be used with a wide range of cloud data store technologies. We implement DAT in AppScale, an open-source implementation of the Google App Engine cloud platform, and use it to evaluate DAT's performance and the performance of a number of popular key-value stores. Navraj Chohan, Chris Bunch, Chandra Krintz, Yoshihide Nomura |
IEEE CLOUD | 3 |
| 2010 | An Evaluation of Distributed Datastores Using the AppScale Cloud PlatformabstractWe present new cloud support that employs a single API the Datastore API from Google App Engine (GAE) to interface to different open source distributed database technologies. We employ this support to "plug in" these technologies to the API so that they can be used by web applications and services without modification. The system facilitates an empirical evaluation and comparison of these disparate systems by web software developers, and reduces the barrier to entry for their use by automating their configuration and deployment. Chris Bunch, Navraj Chohan, Chandra Krintz, Jovan Chohan, Jonathan Kupferman, Puneet Lakhina, Yoshihide Nomura |
IEEE CLOUD | 3 |
| 2010 | Cross-language, type-safe, and transparent object sharing for co-located managed runtimesabstractAs software becomes increasingly complex and difficult to analyze, it is more and more common for developers to use high-level, type-safe, object-oriented (OO) programming languages and to architect systems that comprise multiple components. Different components are often implemented in different programming languages. In state-of-the-art multicomponent, multi-language systems, cross-component communication relies on remote procedure calls (RPC) and message passing. As components are increasingly co-located on the same physical machine to ensure high utilization of multi-core systems, there is a growing potential for using shared memory for cross-language cross-runtime communication. Michal Wegiel, Chandra Krintz |
OOPSLA | 2 |
| 2009 | Dynamic prediction of collection yield for managed runtimesabstractThe growth in complexity of modern systems makes it increasingly difficult to extract high-performance. The software stacks for such systems typically consist of multiple layers and include managed runtime environments (MREs). In this paper, we investigate techniques to improve cooperation between these layers and the hardware to increase the efficacy of automatic memory management in MREs. Michal Wegiel, Chandra Krintz |
ASPLOS | 2 |
| 2009 | As-if-serial exception handling semantics for Java futures
Chandra Krintz |
Sci. Comput. Program. | 2 |
| 2009 | The single-referent collector: Optimizing compaction for the common caseabstractCompactors that move or copy objects need to adjust pointers. In extant compactors, pointer adjustment involves inspecting every pointer in the heap and computing the target address for each pointer. At the same time, in modern Managed Runtime Environments (MREs), only a fraction of pointers in the heap changes during compaction. This is because state-of-the-art MREs do not compact the prefix of the heap that contains few dead objects, allowing gaps between live objects and tolerating small space overhead. We describe the design and implementation of the Single-Referent Collector (SRC), a new compactor that reduces the cost of pointer manipulation by avoiding inspection and adjustment of pointers that do not change. SRC exploits the fact that in modern applications, most live objects have only one incoming pointer. For such objects, SRC stores the address of the referent in the object header. Only objects that move have their referent inspected and adjusted. The remaining pointers in the heap are not processed. SRC uses an overflow table to handle objects with multiple incoming pointers. We investigate a number of standard benchmarks and open-source applications to substantiate key statistical observations that underlie the design of SRC. We implement SRC in the HotSpot JVM as part of a generational collection system and compare it empirically with the Lisp2 compactor. We find that, by decreasing the cost of pointer processing, SRC enables significant reduction in pause times and improves application throughput. Michal Wegiel, Chandra Krintz |
ACM Trans. Archit. Code Optim. | 2 |
| 2008 | The mapping collector: virtual memory support for generational, parallel, and concurrent compactionabstractParallel and concurrent garbage collectors are increasingly employed by managed runtime environments (MREs) to maintain scalability, as multi-core architectures and multi-threaded applications become pervasive. Moreover, state-of-the-art MREs commonly implement compaction to eliminate heap fragmentation and enable fast linear object allocation. Michal Wegiel, Chandra Krintz |
ASPLOS | 2 |
| 2008 | MTM2: Scalable Memory Management for Multi-tasking Managed Runtime Environments
Sunil Soman, Chandra Krintz, Laurent Daynès |
ECOOP | 2 |
| 2008 | Using bandwidth data to make computation offloading decisionsabstractWe present a framework for making computation offloading decisions in computational grid settings in which schedulers determine when to move parts of a computation to more capable resources to improve performance. Such schedulers must predict when an offloaded computation will outperform one that is local by forecasting the local cost (execution time for computing locally) and remote cost (execution time for computing remotely and transmission time for the input/output of the computation to/from the remote system). Typically, this decision amounts to predicting the bandwidth between the local and remote systems to estimate these costs. Our framework unifies such decision models by formulating the problem as a statistical decision problem that can either be treated "classically" or using a Bayesian approach. Using an implementation of this framework, we evaluate the efficacy of a number of different decision strategies (several of which have been employed by previous systems). Our results indicate that a Bayesian approach employing automatic change-point detection when estimating the prior distribution is the best-performing approach. Richard Wolski, Selim Gurun, Chandra Krintz, Daniel Nurmi |
IPDPS | 3 |
| 2008 | XMem: type-safe, transparent, shared memory for cross-runtime communication and coordinationabstractDevelopers commonly build contemporary enterprise applications using type-safe, component-based platforms, such as J2EE, and architect them to comprise multiple tiers, such as a web container, application server, and database engine. Administrators increasingly execute each tier in its own managed runtime environment (MRE) to improve reliability and to manage system complexity through the fault containment and modularity offered by isolated MRE instances. Such isolation, however, necessitates expensive cross-tier communication based on protocols such as object serialization and remote procedure calls. Administrators commonly co-locate communicating MREs on a single host to reduce communication overhead and to better exploit increasing numbers of available processing cores. However, state-of-the-art MREs offer no support for more efficient communication between co-located MREs, while fast inter-process communication mechanisms, such as shared memory, are widely available as a standard operating system service on most modern platforms. Michal Wegiel, Chandra Krintz |
PLDI | 2 |
| 2008 | NWSLite: A general-purpose, nonparametric prediction utility for embedded systemsabstractTime series-based prediction methods have a wide range of uses in embedded systems. Many OS algorithms and applications require accurate prediction of demand and supply of resources. However, configuring prediction algorithms is not easy, since the dynamics of the underlying data requires continuous observation of the prediction error and dynamic adaptation of the parameters to achieve high accuracy. Current prediction methods are either too costly to implement on resource-constrained devices or their parameterization is static, making them inappropriate and inaccurate for a wide range of datasets. This paper presents NWSLite, a prediction utility that addresses these shortcomings on resource-restricted platforms. Selim Gurun, Chandra Krintz, Richard Wolski |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2007 | Call-chain Software Instruction Prefetching in J2EE Server Applications
Priya Nagpurkar, Harold W. Cain, Mauricio J. Serrano, Jong-Deok Choi, Chandra Krintz |
PACT | 5 |
| 2007 | Language and Virtual Machine Support for Efficient Fine-Grained Futures in Java
Chandra Krintz, Priya Nagpurkar |
PACT | 2 |
| 2007 | Isla Vista Heap Sizing: Using Feedback to Avoid PagingabstractManaged runtime environments (MREs) employ garbage collection (GC) for automatic memory management. However, GC induces pressure on the virtual memory (VM) manager, since it may touch pages that are not related to the working set of the application. Paging due to GC can significantly hurt performance, even when the application's working set fits into physical memory. We present a feedback-directed heap resizing mechanism to avoid GC-induced paging, using information from the operating system (OS). We avoid costly GCs when there is physical memory available, and trade off GC for paging when memory is constrained. Our mechanism is simple and uses allocation stall events during GC alone to trigger heap resizing, without user participation or OS kernel modification. Our system enables significant performance improvements when real memory is restricted and similar to, or better performance than, the current state-of-the-art MRE, when memory is unconstrained Chris Grzegorczyk, Sunil Soman, Chandra Krintz, Richard Wolski |
CGO | 3 |
| 2007 | VIProf: Vertically Integrated Full-System Performance ProfilerabstractIn this paper, we present VIProf, a full-system, performance sampling system capable of extracting runtime behavior across an entire software stack. Our long-term goal is to employ VIProf profiles to guide online optimization of programs and their execution environments according to the dynamically changing execution behavior and resource availability. VIProf thus, must be transparent while producing accurate and useful performance profiles. We overview the design and implementation of VIProf and empirically evaluate the system using a popular software stack - one that includes a Linux operating system, a Java virtual machine, and a set of applications. This composition is commonly employed and important for high-end systems such as application and Web servers as well as computational grid services. We show that VIProf introduces little overhead and is able to capture accurate (function-level) full-system performance data that previously required multiple profiles and extensive, manual, and offline post-processing of profile data. Hussam Mousa, Chandra Krintz, Lamia Youseff, Richard Wolski |
IPDPS | 2 |
| 2007 | Application-specific garbage collection
Sunil Soman, Chandra Krintz |
J. Syst. Softw. | 2 |
| 2006 | Online Phase Detection AlgorithmsabstractToday's virtual machines (VMs) dynamically optimize an application as it is executing, often employing optimizations that are specialized for the current execution profile. An online phase detector determines when an executing program is in a stable period of program execution (a phase) or is in transition. A VM using an online phase detector can apply specialized optimizations during a phase or reconsider optimization decisions between phases. Unfortunately, extant approaches to detecting phase behavior rely on either offline profiling, hardware support, or are targeted toward a particular optimization. In this work, we focus on the enabling technology of online phase detection. More specifically, we contribute (a) a novel framework for online phase detection, (b) multiple instantiations of the framework that produce novel online phase detection algorithms, (c) a novel client- and machine-independent baseline methodology for evaluating the accuracy of an online phase detector, (d) a metric to compare online detectors to this baseline, and (e) a detailed empirical evaluation, using Java applications, of the accuracy of the numerous phase detectors. Priya Nagpurkar, Chandra Krintz, Michael Hind, Peter F. Sweeney, V. T. Rajan |
CGO | 2 |
| 2006 | Task-aware garbage collection in a multi-tasking virtual machineabstractA multi-tasking virtual machine (MVM) executes multiple programs in isolation, within a single operating system process. The goal of a MVM is to improve startup time, overall system throughput, and performance, by effective reuse and sharing of system resources across programs (tasks). However, multitasking also mandates a memory management system capable of offering a guarantee of isolation with respect to garbage collection costs, accounting of memory usage, and timely reclamation of heap resources upon task termination.To this end, we investigate and evaluate, novel task-aware extensions to a state-of-the-art MVM garbage collector (GC). Our task-aware GC exploits the generational garbage collection hypothesis, in the context of multiple tasks, to provide performance isolation by maintaining task-private young generations. Task aware GC facilitates concurrent per-task allocation and promotion, and minimizes synchronization and scanning overhead. In addition, we efficiently track per-task heap usage to enable GC-free reclamation upon task termination. Moreover, we couple these techniques with a light-weight synchronization mechanism that enables per-task minor collection, concurrently with allocation by other tasks.We empirically evaluate the efficiency, scalability, and through-put that our task-aware GC system enables. Sunil Soman, Laurent Daynès, Chandra Krintz |
ISMM | 3 |
| 2006 | Phase-based visualization and analysis of Java programs
Priya Nagpurkar, Chandra Krintz |
Sci. Comput. Program. | 2 |
| 2006 | Efficient remote profiling for resource-constrained devicesabstractThe widespread use of ubiquitous, mobile, and continuously connected computing agents has inspired software developers to change the way they test, debug, and optimize software. Users now play an active role in the software evolution cycle by dynamically providing valuable feedback about the execution of a program to developers. Software developers can use this information to isolate bugs in, maintain, and improve the performance of a wide-range of diverse and complex embedded device applications. The collection of such feedback poses a major challenge to systems researchers since it must be performed without degrading a user's experience with, or consuming the severely restricted resources of the mobile device. At the same time, the resource constraints of embedded devices prohibit the use of extant software profiling solutions. To achieve efficient remote profiling of embedded devices, we couple two efficient hardware/software program monitoring techniques: Hybrid Profiling Support(HPS) and Phase-Aware Sampling. HPS efficiently inserts profiling instructions into an executing program using a novel extension to Dynamic-Instruction Stream Editing(DISE). Phase-aware sampling exploits the recurring behavior of programs to identify key opportunities during execution in order to collect profile information (i.e. sample). Our prior work on phase-aware sampling required code duplication to toggle sampling. By guiding low-overhead, hardware-supported sampling according to program phase behavior via HPS, our system is able to collect highly accurate profiles transparently. We evaluate our system assuming a general purpose configuration as well as a popular handheld device configuration. We measure the accuracy and overhead of our techniques and quantify the overhead in terms of computation, communication, and power consumption. We compare our system to random and periodic sampling for a number of widely used performance profile types. Our results indicate that our system significantly reduces the overhead of sampling while maintaining high accuracy. Priya Nagpurkar, Hussam Mousa, Chandra Krintz, Timothy Sherwood |
ACM Trans. Archit. Code Optim. | 3 |
| 2006 | Adaptive On-the-Fly CompressionabstractWe present a system called the adaptive compression environment (ACE) that automatically and transparently applies compression (on-the-fly) to a communication stream to improve network transfer performance. ACE uses a series of estimation techniques to make short-term forecasts of compressed and uncompressed transfer time at the 32 Kb block level. ACE considers underlying networking technology, available resource performance, and data characteristics as part of its estimations to determine which compression algorithm to apply (if any). Our empirical evaluation shows that, on average, ACE improves transfer performance given changing network types and performance characteristics by 8 to 93 percent over using the popular compression techniques that we studied (Bzip, Zlib, LZO, and no compression) alone. Chandra Krintz, Sezgin Sucu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | Phase-Aware Remote ProfilingabstractRecent advances in networking and embedded device technology have made the vision of ubiquitous computing a reality; users can access the Internet's vast offerings anytime and anywhere. Moreover, battery-powered devices such as personal digital assistants and Web-enabled mobile phones have successfully emerged as new access points to the world's digital, infrastructure. This ubiquity offers a new opportunity for software developers: users can now participate in the software development, optimization, and evolution process while they use their software. Such participation requires effective techniques for gathering profile information from remote, resource-constrained devices. Further, these techniques must be unobtrusive and transparent to the user; profiles must be gathered using minimal computation, communication, and power. Toward this end, we present a flexible hardware-software scheme for efficient remote profiling. We rely on the extraction of meta information from executing programs in the form of phases, and then use this information to guide intelligent online sampling and to manage the communication of those samples. Our results indicate that phase-based remote profiling can reduce the communication, computation, and energy consumption overheads by 50-75% over random and periodic sampling. Priya Nagpurkar, Chandra Krintz, Timothy Sherwood |
CGO | 2 |
| 2005 | AutoDVS: an automatic, general-purpose, dynamic clock scheduling system for hand-held devicesabstractWe present AutoDVS, a dynamic voltage scaling (DVS) system for hand-held computers. Unlike extant DVS systems, AutoDVS distinguishes common, course-grain, program behavior and couples forecasting techniques to make accurate predictions of future behavior. AutoDVS uses these predictions in combination to guide dynamic voltage scaling. AutoDVS estimates periods of user interactivity, user non-interactivity (think time), and computation per-program and system wide to ensure quality of service while reducing energy consumption.We describe our implementation of AutoDVS which consists of a set light-weight, Linux, kernel modules and user library routines for the iPAQ hand-held computer. We evaluate AutoDVS using real user workloads of iPAQ software that consist of interactive and soft-real time tasks executing alone and concurrently. Our results indicate that AutoDVS decreases energy consumption significantly without negatively impacting user perception of system performance. Selim Gurun, Chandra Krintz |
EMSOFT | 2 |
| 2005 | The design, implementation, and evaluation of adaptive code unloading for resource-constrained devicesabstractJava Virtual Machines (JVMs) for resource-constrained devices, e.g., hand-helds and cell phones, commonly employ interpretation for program translation. However, compilers are able to produce significantly better code quality, and, hence, use device resources more efficiently than interpreters, since compilers can consider large sections of code concurrently and exploit optimization opportunities. Moreover, compilation-based systems store code for reuse by future invocations obviating the redundant computation required for reinterpretation of repeatedly executed code.However, code storage required for compilation can increase the memory footprint of the virtual machine (VM) significantly. As a result, for devices with limited memory resources, this additional code storage may preclude some programs from executing, significantly increase memory management overhead, and substantially reduce the amount of memory available for use by the application.To address the limitations of native code storage, we present the design, implementation, and empirical evaluation of a compiled-code management system that can be integrated into any compilation-based JVM. The system unloads compiled code to reduce the memory footprint of the VM. It does so by dynamically identifying and unloading dead or infrequently used code; if the code is later reused, it is recompiled by the system. As such, our system adaptively trades off memory footprint and its associated memory management costs, with recompilation overhead. Our empirical evaluation shows that our code management system significantly reduces the memory requirements of a compile-only JVM, while maintaining the performance benefits enabled by compilation.We investigate a number of implementation alternatives that use dynamic program behavior and system resource availability to determine when to unload as well as what code to unload. From our empirical evaluation of these alternatives, we identify a set of strategies that enable significant reductions in the memory overhead required for application code. Our system reduces code size by 36--62%, on average, which translates into significant execution-time benefits for the benchmarks and JVM configurations that we studied. Chandra Krintz |
ACM Trans. Archit. Code Optim. | 2 |
| 2004 | Application-level prediction of battery dissipationabstractMobile, battery-powered devices such as personal digital assistants and web-enabled mobile phones have successfully emerged as new access points to the world's digital infrastructure. However, the growing gap between device capabilities and battery technology requires novel techniques that extend battery life. Key to the success of such techniques, is our ability to accurately predict the power consumption of a program.In this paper, we investigate the degree to which battery dissipation induced by program execution can be measured by applica-tion-level software tools and predicted by a compiler and runtime system. We present a novel technique with which we can accurately estimate whole-program power-consumption for an arbitrary program by composing battery dissipation rates of benchmarks. We empirically evaluate our technique using an iPAQ hand-held device and a number of MiBench and other programs. Chandra Krintz, Ye Wen, Richard Wolski |
ISLPED | 1 |
| 2004 | Dynamic selection of application-specific garbage collectorsabstractMuch prior work has shown that the performance enabled by garbage collection (GC) systems is highly dependent upon the behavior of the application as well as on the available resources. That is, no single GC enables the best performance for all programs and all heap sizes. To address this limitation, we present the design, implementation, and empirical evaluation of a novel Java Virtual Machine (JVM) extension that facilitates dynamic switching between a number of very different and popular garbage collectors. We also show how to exploit this functionality using annotation-guided GC selection and evaluate the system using a large number of benchmarks. In addition, we implement and evaluate a simple heuristic to investigate the efficacy of switching automatically. Our results show that, on average, our annotation-guided system introduces less than 4% overhead and improves performance by 24% over the worst-performing GC (across heap sizes) and by 7% over always using the popular Generational/Mark-Sweep hybrid. Sunil Soman, Chandra Krintz, David F. Bacon |
ISMM | 2 |
| 2004 | Adaptive code unloading for resource-constrained JVMsabstractCompile-only JVMs for resource-constrained embedded systems have the potential for using device resources more efficiently than interpreter-only systems since compilers can produce significantly higher quality code and code can be stored and reused for future invocations. However, this additional storage requirement for reuse of native code bodies, introduces memory overhead not imposed in interpreter-based systems.In this paper, we present a Java Virtual Machine (JVM) extension for adaptive code unloading that significantly reduces the memory requirements imposed by a compile-only JVM. The extension features an unloader that uses execution behavior to adaptively determine when to unload as well as what code to unload. We implement and empirically identify a set of unloading strategies that enable significant code size reduction (43%-61%). This reduction translates into significant execution time benefits for the benchmarks and JVM configurations that we studied. As such, by using adaptive code unloading, we make compile-only JVMs for embedded devices more feasible. Chandra Krintz |
LCTES | 2 |
| 2004 | NWSLite: A Light-Weight Prediction Utility for Mobile DevicesabstractComputation off-loading, i.e., remote execution, has been shown to be effective for extending the computational power and battery life of resource-restricted devices, e.g., hand-held, wearable, and pervasive computers. Remote execution systems must predict the cost of executing both locally and remotely to determine when off-loading will be most beneficial. These costs however, are dependent upon the execution behavior of the task being considered and the highly-variable performance of the underlying resources, e.g., CPU (local and remote), bandwidth, and network latency. As such, remote execution systems must employ sophisticated, prediction techniques that accurately guide computation off-loading. Moreover, these techniques must be efficient, i.e., they cannot consume significant resources, e.g., energy, execution time, etc., since they are performed on the mobile device.In this paper, we present NWSLite, a computationally efficient, highly accurate prediction utility for mobile devices. NWSLite is an extension to the Network Weather Service (NWS), a dynamic forecasting toolkit for adaptive scheduling of high-performance Computational Grid applications. We significantly scaled down the NWS to reduce its resource consumption yet still achieve accuracy that exceeds that of extant remote execution prediction methods. We empirically analyze and compare both the prediction accuracy and the cost of NWSLite and a number of different forecasting methods from existing remote execution systems. We evaluate the efficacy of the different methods using a wide range of mobile applications and resources. Selim Gurun, Chandra Krintz, Richard Wolski |
MobiSys | 2 |
| 2003 | Coupling On-Line and Off-Line Profile Information to Improve Program PerformanceabstractIn this paper we describe a novel execution environment for Java programs that substantially improves execution performance by incorporating both on-line and off-line profile information to guide dynamic optimization. By using both types of profile collection techniques, we are able to exploit the strengths of each constituent approach: profile accuracy and low overhead. Such coupling also reduces the negative impact of these approaches when each is used in isolation. On-line profiling introduces overhead for dynamic instrumentation, measurement, and decision making. Off-line profile information can be inaccurate when program inputs for execution and optimization differ from those used for profiling. To combat these drawbacks and to achieve the benefits from both online and off-line profiling, we developed a dynamic compilation system (based on JikesRVM) that makes use of both. As a result, we are able improve Java program performance by 9% on average, for the programs studied. Chandra Krintz |
CGO | 1 |
| 2003 | Detecting Malicious Java Code Using Virtual Machine Auditing
Sunil Soman, Chandra Krintz, Giovanni Vigna |
USENIX Security Symposium | 2 |
| 2001 | NwsAlarm: A Tool for Accurately Detecting Resource Performance DegradationabstractEnd-users of high-performance computing resources have come to expect that consistent levels of performance be delivered to their applications. The advancement of the computational grid enables the seamless use of a multitude of computing resources by these users. The combination of these developments has generated a need for users to monitor the end-to-end-performance available to an application. In addition, tools are needed to alert users of degradation in expected performance. We present the NwsAlarm, a Java-based utility that enables users to monitor performance levels of any resource being monitored by the Network Weather Service. The NwsAlarm is invoked by a user without special privileges with a simple click on a Web page link. More importantly the NwsAlarm allows any user of the NwsAlarm to register and set expected performance levels. When performance levels fall below these thresholds, the registered administrators are immediately notified via email. The NwsAlarm uses prediction of performance measurements to filter false alarm values. We exemplify the importance of and accuracy achieved by the NwsAlarm with real examples of performance degradation caused by routing table changes and loss of service on the Abilene, Internet-2 research network used for experimentation with evolving Grid software technology. On average, 92% fewer false alarms are raised by the NwsAlarm than if raw measurements are used. Chandra Krintz, Richard Wolski |
CCGRID | 1 |
| 2001 | Reducing Delay with Dynamic Selection of Compression FormatsabstractInternet computing is facilitated by a remote execution methodology in which programs transfer to a destination for execution. Since the transfer time can substantially degrade the performance of remotely executed (mobile) programs, file compression is used to reduce the amount of data that is transferred. Compression techniques however, must trade off compression ratio for decompression time, due to the algorithmic complexity of the former, since the latter is performed at run-time in this environment. In this paper, we define the total delay as the time for both the transfer and the decompression of a compressed file. To minimize the total delay, a mobile program should be compressed in the best format for minimizing the delay. Since both the transfer time and the decompression time are dependent upon the current underlying resource performance, selection of the "best" format varies and no one compression format minimizes the total delay for all resource performance characteristics. We present a system called Dynamic Compression Format Selection (DCFS) for the automatic and dynamic selection of competitive compression formats based on the predicted values of future resource performance. Our results show that DCFS reduces the total delay imposed by the compressed transfer of Java archives (.jar files) by 52% on average for the networks, compression techniques and benchmarks studied. Chandra Krintz, Brad Calder |
HPDC | 1 |
| 2001 | Using Annotation to Reduce Dynamic Optimization TimeabstractDynamic compilation and optimization are widely used in heterogenous computing environments, in which an intermediate form of the code is compiled to native code during execution. An important trade off exists between the amount of time spent dynamically optimizing the program and the running time of the program. The time to perform dynamic optimizations can cause significant delays during execution and also prohibit performance gains that result from more complex optimization. Chandra Krintz, Brad Calder |
PLDI | 1 |
| 2001 | Using JavaNws to compare C and Java TCP-Socket performanceabstractAbstract As research and implementation continue to facilitate high‐performance computing in Java, applications can benefit from resource management and prediction tools. In this work, we present such a tool for network round‐trip time and bandwidth between a user's desktop and any machine running a Web server (this assumes that the user's browser is capable of supporting Java 1.1 and above). JavaNws is a Java implementation and extension of a powerful subset of the Network Weather Service (NWS), a performance prediction toolkit that dynamically characterizes and forecasts the performance available to an application. However, due to the Java language implementation and functionality (portability, security, etc.), it is unclear whether a Java program is able to measure and predict the network performance experienced by C‐applications with the same accuracy as an equivalent C program. We provide a quantitative equivalence study of the Java and C TCP‐socket interface and show that the data collected by the JavaNws is as predictable as that collected by the NWS (using C). Copyright © 2001 John Wiley & Sons, Ltd. Chandra Krintz, Richard Wolski |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | Reducing the overhead of dynamic compilationabstractAbstract The execution model for mobile, dynamically‐linked, object‐oriented programs has evolved from fast interpretation to a mix of interpreted and dynamically compiled execution. The primary motivation for dynamic compilation is that compiled code executes significantly faster than interpreted code. However, dynamic compilation, which is performed while the application is running, introduces execution delay. In this paper we present two dynamic compilation techniques that enable high performance execution while reducing the effect of this compilation overhead. These techniques can be classified as (1) decreasing the amount of compilation performed, and (2) overlapping compilation with execution. We first present and evaluate lazy compilation , an approach used in most dynamic compilation systems in which individual methods are compiled on‐demand upon their first invocation. This is in contrast to eager compilation , in which all methods in a class are compiled when a new class is loaded. In this work, we describe our experience with eager compilation, as well as the implementation and transition to lazy compilation. We empirically detail the effectiveness of this decision. Our experimental results using the SpecJVM Java benchmarks and the Jalapeño JVM show that, compared to eager compilation, lazy compilation results in 57% fewer methods being compiled and reductions in total time of 14 to 26%. Total time in this context is compilation plus execution time. Next, we present profile‐driven, background compilation, a technique that augments lazy compilation by using idle cycles in multiprocessor systems to overlap compilation with application execution. With this approach, compilation occurs on a thread separate from that of application threads so as to reduce intermittent, and possibly substantial, delay in execution. Profile information is used to prioritize methods as candidates for background compilation. Methods are compiled according to this priority scheme so that performance‐critical methods are invoked using optimized code as soon as possible. Our results indicate that background compilation can achieve the performance of off‐line compiled applications and masks almost all compilation overhead. We show significant reductions in total time of 14 to 71% over lazy compilation. Copyright © 2001 John Wiley & Sons, Ltd. Chandra Krintz, David Grove, Vivek Sarkar, Brad Calder |
Softw. Pract. Exp. | 1 |
| 1999 | Reducing Transfer Delay Using Java Class File Splitting and PrefetchingabstractThe proliferation of the Internet is fueling the development of mobile computing environments in which mobile code is executed on remote sites. In such environments, the end user must often wait while the mobile program is transferred from the server to the client where it executes. This downloading can create significant delays, hurting the interactive experience of users. Chandra Krintz, Brad Calder, Urs Hölzle |
OOPSLA | 1 |
| 1999 | Running EveryWare on the Computational GridabstractThe Computational Grid [10] has recently been proposed for the implementation of high-performance applications using widely dispersed computational resources. The goal of a Computational Grid is to aggregate ensembles of shared, heterogeneous, and distributed resources (potentially controlled by separate organizations) to provide computational "power" to an application program. In this paper, we provide a toolkit for the development of Grid applications. The toolkit, called EveryWare, enables an application to draw computational power transparently from the Grid. The toolkit consists of a portable set of processes and libraries that can be incorporated into an application so that a wide variety of dynamically changing distributed infrastructures and resources can be used together to achieve supercomputer-like performance. We provide our experiences gained while building the EveryWare toolkit prototype and the first true Grid application. 1 Introduction Increasingly, the high-perform... Richard Wolski, John Brevik, Chandra Krintz, Graziano Obertelli, Neil Spring, Alan Su 0001 |
SC | 3 |
| 1998 | Cache-Conscious Data PlacementabstractAs the gap between memory and processor speeds continues to widen, cache eficiency is an increasingly important component of processor performance. Compiler techniques have been used to improve instruction cache pet$ormance by mapping code with temporal locality to different cache blocks in the virtual address space eliminating cache conflicts. These code placement techniques can be applied directly to the problem of placing data for improved data cache pedormance.In this paper we present a general framework for Cache Conscious Data Placement. This is a compiler directed approach that creates an address placement for the stack (local variables), global variables, heap objects, and constants in order to reduce data cache misses. The placement of data objects is guided by a temporal relationship graph between objects generated via profiling. Our results show that profile driven data placement significantly reduces the data miss rate by 24% on average. Brad Calder, Chandra Krintz, Simmi John, Todd M. Austin |
ASPLOS | 2 |
| 1998 | Overlapping Execution with Transfer Using Non-Strict Execution for Mobile ProgramsabstractIn order to execute a program on a remote computer, it mustfirst be transferred over a network. This transmission incurs the over-head of network latency before execution can begin. This latency can vary greatly depending upon the size of the program., where it is located (e.g., on a local network or across the Internet), and the bandwidth available to retrieve the program. Existing technologies, like Java, require that a jle be filly transferred before it can start executing. For large files and low bandwidth lines, this delay can be significant.In this paper we propose and evaluate a non-strict form of mobile program execution. A mobile program is any program that is transferred to a different machine and executed. The goal of nonstrict execution is to overlap execution with transfer; allowing the program to start executing as soon as possible. Non-strict execution allows a procedure in the program to start executing as soon as its code and data have transferred. To enable this technology, we examine several techniques for rearranging procedures and reorganizing the data inside Java classjles. Our results show that nonstrict execution decreases the initial transfer delay between 31% and 56% on average, with an average reduction in overall execution time between 25% and 40%. Chandra Krintz, Brad Calder, Han Bok Lee, Benjamin G. Zorn |
ASPLOS | 1 |