VLDB 2026 Research / reviewers in the wild / expert
Ada Gavrilovska
dblp:76/3229
· DBLP profile ↗
84ranked-venue papers
9as first author
33since 2021 · last 2026
0000-0003-4199-2512ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 6 first-author · 26 since 2021Computer networks · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Security and privacy · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL InterconnectsabstractCXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline. Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska |
ICS | 6 |
| 2026 | AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory Systems
Suyeon Lee, Kangkyu Park, Kwangsik Shin, Ada Gavrilovska |
ISCA | 4 |
| 2026 | Stimpack: An Adaptive Rendering Optimization System for Scalable Cloud Gaming
Jin Heo, Vic Wang, Ketan Bhardwaj, Ada Gavrilovska |
NSDI | 4 |
| 2025 | Grudon: A System for Deploying Graph Workloads on Disaggregated Architectures with Near-Data ProcessingabstractEmerging memory disaggregation and near-data processing (NDP) technologies have shown promise in addressing performance, scaling and efficiency challenges for data-intensive workloads such as graph analytics. However, they expose new tradeoffs related to data movement, runtime, and cost. This paper introduces Grudon, a novel system for deploying graph workloads on Disaggregated NDP (DiNDP) platforms. Grudon combines the benefits of memory disaggregation and NDP to tackle bandwidth limitations, improve resource flexibility, and reduce energy demands to improve overall workload performance. To achieve this, Grudon provides support for dynamic configuration of graph workload deployments on DiNDP platforms via new control plane functionality that enables lightweight evaluation of the tradeoff space and dynamic reconfiguration of operations executed on the different DiNDP elements. Evaluation on an emulated DiNDP testbed demonstrates that Grudon achieves a geomean speedup of 2.12× while reducing energy consumption by 3.3× compared to state-of-the-art distributed runtimes. These results show the potential to unlock the full capabilities of DiNDP platforms for efficient distributed graph analytics. Vishal Rao, Nikhil Ram Shashidhar, Suyeon Lee, Ada Gavrilovska |
HPDC | 4 |
| 2025 | Introduction to the Special Section on USENIX OSDI 2024
Ada Gavrilovska, Douglas B. Terry |
ACM Trans. Storage | 1 |
| 2024 | Krios: Scheduling Abstractions and Mechanisms for Enabling a LEO Compute CloudabstractLow Earth Orbit (LEO) satellites are an important facet of global connectivity providing high speed Internet, cellular, IoT connectivity and so on. Combined with the rich resource availability on each satellite, LEO satellites represent a new, emerging cloud frontier - the LEO Compute Cloud. However, satellite mobility introduces non-trivial challenges when orchestrating applications for a LEO compute cloud, making it harder to deploy applications without increasing the latency and bandwidth costs. In this paper, we identify the concrete challenges in using state-of-the-art terrestrial orchestrators for a LEO compute cloud. We present Krios - a LEO compute cloud orchestration system that hides the complexities introduced by satellite mobility and enables a practical LEO compute cloud. The design of Krios is centered around a novel LEO zones abstraction that allows application providers to specify where their applications should be available. Krios provides crucial system support to enable the LEO zones abstraction, ensuring uninterrupted availability of applications in LEO zones via proactive and predictive application handovers. Our experimental evaluation of Krios with representative applications demonstrates a practical and efficient LEO compute cloud, without requiring any disruptive changes in applications and with modest system overheads. With Krios, LEO orchestration requires just ~1 application instance at a time to maintain the same availability as what prior work achieves by deploying application instances on all satellites or by performing 6-10 times more frequent expensive handovers. Vaibhav Bhosale, Ada Gavrilovska, Ketan Bhardwaj |
SoCC | 2 |
| 2024 | Poster: Adapting XR Perception Serving for Edge Server ScalabilityabstractOffloading perception tasks to an edge server enhances extended reality (XR) experiences on resource-constrained mobile devices. However, when serving multiple users, an edge server may face resource contention and not be able to meet the service level objectives (SLOs), leading to degraded user experiences. To improve server scalability, we propose a system that adaptively schedules the inference executions of object detection models based on estimated effectiveness. The perception effectiveness is estimated by the inter-frame similarity and the object distance. We present the initial design of the system and preliminary results demonstrating its feasibility. We describe the next steps to further develop our system. Jin Heo, Ada Gavrilovska |
SEC | 2 |
| 2024 | GT-Craft: A Framework for Fast Prototyping Geospatial-Based Digital Twins in Unity 3DabstractA digital twin presents promising opportunities and potential benefits for various industrial use cases by enabling simulation and prediction on the virtual representation of the real-world environment. However, the implementation and maintenance costs for the digital twin are prohibitively high, restricting its widespread adoption. To address this issue, we present a framework, GT-Craft, which enables fast prototyping the geospatial-based digital twin at scale. GT-Craft automates the generation of the digital twin by using the streamed geospatial data and the semantic information extracted from deep neural network (DNN) models. As GT-Craft generates digital twins on the Unity game engine, the Unity-based simulators and game applications can seamlessly use the digital twins generated by GT-Craft. The presented framework is compatible with non-Unity-based applications and existing 3D software and simulation tools, e.g., Blender, Apple Reality Composer, and NVIDIA Omniverse, as it supports exporting the generated digital twin in the universal scene description (USD) format, which is an emerging industrial open standard for exchanging and editing 3D contents. Jin Heo, Thomas David Novlan, Salam Akoum, Ada Gavrilovska |
SEC | 4 |
| 2024 | Colibri: Efficient Collection of Fine-Grained Resource Metrics Necessary for Mobile Edge ComputingabstractEffective provisioning and resource management in edge environments are critical for ensuring infrastructure efficiency while providing latency-critical service-level objectives (SLOs). Realizing this requires low-overhead aggregation of fine-grained information regarding workloads' resource demands. Unfortunately, we show that this cannot be adequately achieved with existing solutions used for cloud technologies such as Kubernetes and containers, which are prevalent in edge systems. We propose Colibri, a lightweight and flexible monitoring system for edge computing that characterizes containers across CPU, memory, and network resource usage patterns at millisecond granularity. Colibri can be dispatched dynamically as needed and enables accurate characterization of workload resource usage. We demonstrate experimentally that using Colibri can significantly reduce SLO violations caused when relying on existing tools, while saving resources for representative edge workloads. Colibri provides this while consuming only 2% of the resources used by existing cloud monitoring tools when they operate at the same query granularity, making it an efficient and effective solution for edge computing environments. Ke-Jou Hsu, Ketan Bhardwaj, Ada Gavrilovska |
SEC | 3 |
| 2024 | Efficient Cross-Frequency Beam Prediction in 6G Wireless Using Time Series DataabstractNext generation 6G wireless envisions a much higher data rate and a lower latency compared to 5G wireless networks. Directional antennas with narrow beams across high mmWave frequencies hold the key to achieving such high data rates. However, Beam Management (BM), which is the process of finding appropriate transmit and receive beams, offers significant challenges. Dynamic channel variation, user mobility, and narrow beams in high frequency mmWave channels further complicate these challenges. Efficient Machine Learning (ML) strategies can be used to alleviate this overhead. In spite of their underlying differences, sub-6 GHz and high frequency mmWave channels share some similarities in array geometry, number of paths, and surrounding environment. As sub-6 GHz channel characteristics are relatively easier to acquire and learn, we introduce a new machine learning framework, using transformer and LSTM to learn sub-6 GHz channel information over time for efficient beam prediction across high frequency mmWave channels. System Level simulation results point out that when using time-series data-based learning of beam patterns with transformers or LSTMs in sub-6 GHz channels, our proposed scheme achieves up to 99.5% top-5 beam prediction accuracy while reducing the BM overhead by over 50% compared to existing work. Vaibhav Bhosale, Navrati Saxena, Ketan Bhardwaj, Ada Gavrilovska, Abhishek Roy 0001 |
ISNCC | 4 |
| 2024 | Snail: Secure Single Iteration LocalizationabstractLocalization is a computer vision task by which the position and orientation of a camera is determined from an image and environmental map. We propose a method for performing localization in a privacy preserving manner supporting two scenarios: first, when the image and map are held by a client who wishes to offload localization to untrusted third parties, and second, when the image and map are held separately by untrusting parties. Privacy preserving localization is necessary when the image and map are confidential, and offloading conserves on-device power and frees resources for other tasks. To accomplish this we integrate existing localization methods and secure multi-party computation (MPC), specifically garbled circuits, yielding proof-based security guarantees in contrast to existing obfuscation-based approaches which recent related work has shown vulnerable. We present two approaches to localization, a baseline data-oblivious adaptation of localization suitable for garbled circuits and our novel Single Iteration Localization. Our technique improves overall performance while maintaining confidentiality of the input image, map, and output pose at the expense of increased communication rounds but reduced computation and communication required per round. Single Iteration Localization is over two orders of magnitude faster than a straightforward application of garbled circuits to localization enabling real-world usage in Turbo the Snail, the first robot to offload localization without revealing input images, environmental map, position, or orientation to offload servers. James Choncholas, Pujith Kachana, André Mateus 0001, Gregoire Phillips, Ada Gavrilovska |
Proc. Priv. Enhancing Technol. | 5 |
| 2023 | Flame: Simplifying Topology Extension in Federated LearningabstractDistributed machine learning approaches, including a broad class of federated learning (FL) techniques, present a number of benefits when deploying machine learning applications over widely distributed infrastructures. The benefits are highly dependent on the details of the underlying machine learning topology, which specifies the functionality executed by the participating nodes, their dependencies and interconnections. Current systems lack the flexibility and extensibility necessary to customize the topology of a machine learning deployment. We present Flame, a new system that provides flexibility of the topology configuration of distributed FL applications around the specifics of a particular deployment context, and is easily extensible to support new FL architectures. Flame achieves this via a new high-level abstraction Topology Abstraction Graphs (TAGs). TAGs decouple the ML application logic from the underlying deployment details, making it possible to specialize the application deployment with reduced development effort. Flame is released as an open source project, and its flexibility and extensibility support a variety of topologies and mechanisms, and can facilitate the development of new FL methodologies. Harshit Daga, Jaemin Shin 0005, Dhruv Garg, Ada Gavrilovska, Myungjin Lee, Ramana Rao Kompella |
SoCC | 4 |
| 2023 | Pocket: ML Serving from the EdgeabstractOne of the major challenges in serving ML applications is the resource pressure introduced by the underlying ML frameworks. This becomes a bigger problem at resource-constrained, multi-tenant edge server locations, where it is necessary to scale to a larger number of clients with a fixed resource envelope. Naive approaches which simply minimize the resource budget allocation of each application result in performance degradation that voids the benefits expected from operating at the edge. Misun Park, Ketan Bhardwaj, Ada Gavrilovska |
EuroSys | 3 |
| 2023 | In-Network Compression for Accelerating IoT Analytics at ScaleabstractTo enable the Internet of Things (IoT) to scale at the level of next generation smart cities and grids, there is a need for cost-effective infrastructure for hosting IoT analytics applications. Offload and acceleration via SmartNICs have been shown to provide benefits to these workloads. However, even with offload, long-term analysis on IoT data still needs to operate on massive number of device updates, often in the form of small messages. Despite offloading, the ingestion of these updates continues to present server bottlenecks. In this paper, we present domain-specific compression and batching engines, that leverage the unique properties of IoT messages to reduce the load on analytics servers and improve their scalability. Using a prototype system based on the InnovaFlex programmable SmartNICs, and several representative IoT benchmarks, we demonstrate that the combination of these techniques achieves up to 14.5× improvement in sustained throughput rates compared to a system without SmartNIC offload, and up to 7× improvement over existing offload approaches. Rafael Oliveira 0011, Ada Gavrilovska |
HOTI | 2 |
| 2023 | Angler: Dark Pool Resource AllocationabstractDemand for distributed computational infrastructure is growing in order to offer low latency connections to end users. The fragmenting infrastructure complicates the resource allocation process. As the number of infrastructure providers grows, points of presence are resource constrained compared to the cloud, they have diverse availability profiles, and diverse connectivity properties. Existing resource allocation approaches require providers share intimate details about their infrastructure to support the placement process, or rely on third party aggregators. Such solutions introduce strong assumptions of trust and collaboration. In this work we present Angler, the first system to allocate resources from dark pools, meaning the capacity and requests of the distributed pool of resources are unknown. Angler leverages cryptographic protocols for secure function evaluation, namely the WRK secure multiparty computation (MPC) protocol [76]. While MPC protocols can have large overheads compared to plaintext function evaluation, an end-to-end approach to the system design subverts the expensive overheads. Specifically, Angler combines a tuned implementation of a maliciously secure MPC protocol, a tailored distributed hash table, and a systematic effort to make the best allocation decision within a response time envelope. Angler is only 2x slower than resource allocation with no privacy when arbitrating among 8 providers, taking less than a second. James Choncholas, Ketan Bhardwaj, Vladimir Kolesnikov, Ada Gavrilovska |
SEC | 4 |
| 2023 | Demo: Privacy-Preserving Localization for Edge-Assisted Robotic NavigationabstractLocalization is a common task in computer vision by which the position and orientation of a camera is determined from an image and 3D map of the environment. We propose a demonstration of securely performing localization in a privacy preserving manner. This allows localization so that a lightweight client device, such as a mobile robot, may offload the computation to an untrusted server without revealing anything about their data. To accomplish this goal, we combine existing localization methods with secure multiparty computation (MPC), specifically garbled circuits. As such, the security guarantees of this work are simulation-based in contrast to existing obfuscation-based approaches to pose estimation for which privacy is inversely proportional to input size. We propose two approaches, a baseline data-oblivious adaptation of localization suitable for MPC and Single Iteration Localization which runs each localization step individually, improving performance at the cost of round complexity while maintaining confidentiality of the input image, map, and output pose. Single Iteration Localization is over two orders of magnitude faster than the data-oblivious approach enabling real-world usage in Turbo the Snail, the first robot to offload localization without revealing input images, environmental map, position, or orientation to offload servers. James Choncholas, Pujith Kachana, André Mateus 0001, Gregoire Phillips, Ada Gavrilovska |
SEC | 5 |
| 2023 | PinIt: Influencing OS Scheduling via Compiler-Induced AffinitiesabstractIn multi-core machines, applications execute in a complex-co-execution environment in which the number of concurrently executing applications typically exceed the number of available cores. In order to fairly and efficiently utilize cores, modern operating systems (OS) such as Linux migrate threads between cores during execution. Although such thread migrations alleviate the problem of stalling and load balancing yielding better core utilization, they also tend to destroy data locality, resulting in fewer cache hits, TLB hits, and thus performance loss for the group of applications collectively. This problem is especially severe in embedded servers which execute media and vision applications that exhibit high data locality. One one hand, mitigating this problem across a group of applications based on OS only solution is infeasible since OS treats applications as blackboxes and has no knowledge of its locality and other behavior. On the other hand, to-date, compiler optimization have focused on analysis, transformations and performance enhancement of applications in isolation ignoring the problem of optimizing performance for applications as a group. This is because of the infeasibility of global-compiler analysis across applications as well as due to the dynamic nature of inter-application interactions which is statically unknown. Girish Mururu, Kangqi Ni, Ada Gavrilovska, Santosh Pande |
LCTES | 3 |
| 2023 | FleXR: A System Enabling Flexibly Distributed Extended RealityabstractExtended reality (XR) applications require computationally demanding functionalities with low end-to-end latency and high throughput. To enable XR on commodity devices, a number of distributed systems solutions enable offloading of XR workloads on remote servers. However, they make a priori decisions regarding the offloaded functionalities based on assumptions about operating factors, and their benefits are restricted to specific deployment contexts. To realize the benefits of offloading in various distributed environments, we present a distributed stream processing system, FleXR, which is specialized for real-time and interactive workloads and enables flexible distributions of XR functionalities. In building FleXR, we identified and resolved several issues of presenting XR functionalities as distributed pipelines. FleXR provides a framework for flexible distribution of XR pipelines while streamlining development and deployment phases. We evaluate FleXR with three XR use cases in four different distribution scenarios. In the results, the best-case distribution scenario shows up to 50% less end-to-end latency and 3.9x pipeline throughput compared to alternatives. Jin Heo, Ketan Bhardwaj, Ada Gavrilovska |
MMSys | 3 |
| 2023 | A Characterization of Route Variability in LEO Satellite Networks
Vaibhav Bhosale, Ahmed Saeed 0001, Ketan Bhardwaj, Ada Gavrilovska |
PAM | 4 |
| 2023 | Beacons: An End-to-End Compiler Framework for Predicting and Utilizing Dynamic Loop CharacteristicsabstractEfficient management of shared resources is a critical problem in high-performance computing (HPC) environments. Existing workload management systems often promote non-sharing of resources among different co-executing applications to achieve performance isolation. Such schemes lead to poor resource utilization and suboptimal process throughput, adversely affecting user productivity. Tackling this problem in a scalable fashion is extremely challenging, since it requires the workload scheduler to possess an in-depth knowledge about various application resource requirements and runtime phases at fine granularities within individual applications. In this work, we show that applications’ resource requirements and execution phase behaviour can be captured in a scalable and lightweight manner at runtime by estimating important program artifacts termed as “ dynamic loop characteristics ”. Specifically, we propose a solution to the problem of efficient workload scheduling by designing a compiler and runtime cooperative framework that leverages novel loop-based compiler analysis for resource allocation . We present Beacons Framework , an end-to-end compiler and scheduling framework, that estimates dynamic loop characteristics, encapsulates them in compiler-instrumented beacons in an application, and broadcasts them during application runtime, for proactive workload scheduling. We focus on estimating four important loop characteristics : loop trip-count , loop timing , loop memory footprint , and loop data-reuse behaviour , through a combination of compiler analysis and machine learning. The novelty of the Beacons Framework also lies in its ability to tackle irregular loops that exhibit complex control flow with indeterminate loop bounds involving structure fields, aliased variables and function calls , which are highly prevalent in modern workloads. At the backend, Beacons Framework entails a proactive workload scheduler that leverages the runtime information to orchestrate aggressive process co-locations, for maximizing resource concurrency, without causing cache thrashing . Our results show that Beacons Framework can predict different loop characteristics with an accuracy of 85% to 95% on average, and the proactive scheduler obtains an average throughput improvement of 1.9x (up to 3.2x ) over the state-of-the-art schedulers on an Amazon Graviton2 machine on consolidated workloads involving 1000-10000 co-executing processes, across 51 benchmarks. Girish Mururu, Sharjeel Khan, Bodhisatwa Chatterjee, Chao Chen 0024, Chris Porter, Ada Gavrilovska, Santosh Pande |
Proc. ACM Program. Lang. | 6 |
| 2023 | CLUE: Systems Support for Knowledge Transfer in Collaborative Learning With Neural NetsabstractFor highly distributed environments such as edge computing, collaborative learning approaches eschew the dependence on a global, shared model, in favor of models tailored for each location. Creating tailored models for individual learning contexts reduces the amount of data transfer, while collaboration among peers provides acceptable model performance. Collaboration assumes, however, the availability of knowledge transfer mechanisms, which are not trivial for deep learning models where knowledge isn't easily attributed to precise model slices. We present CLUE – a framework that facilitates knowledge transfer for neural networks. CLUE provides new system support for dynamically extracting significant parameters from a helper node's neural network, and uses this with a multi-model boosting-based approach to improve the predictive performance of the target node. The evaluation of CLUE with different PyTorch and TensorFlow neural network models demonstrates that its knowledge transfer mechanism improves by up to$3.5\times$how quickly a model adapts to changes, compared to learning in isolation, while affording up to several magnitudes reduction in data movement costs compared to federated learning. Harshit Daga, Aastha Agrawal, Ada Gavrilovska |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Coeus: Clustering (A)like Patterns for Practical Machine Intelligent Hybrid Memory ManagementabstractEmerging workloads benefit from massive memory capacities provided by hybrid memory platforms. Recent system-level hybrid memory management solutions integrate machine learning methods to better predict complex data access behaviors. Given the substantial associated learning overheads, such solutions train parallel recurrent neural networks to learn the access patterns at the granularity of a page for a carefully selected page subset. Our observation reveals that the size of this subset varies immensely across workload classes, sizes and patterns. Increasing the granularity at the level of a page group will help reduce the aggregate learning overheads. Yet, unsupervised machine learning clustering methods are not practical to use in this context. Instead, this paper builds Coeus - a page grouping mechanism for machine learning-based hybrid memory management. Coeus is simple, robust and efficient. Coeus leverages data reuse insights to fine-tune the granularity at which patterns are interpreted by the system. As a result, Coeus creates large clusters of pages that share the same access behavior, in a practical way. Coeus reduces by almost 3x the associated learning overheads. In addition, Coeus achieves 3x higher application performance, by the combined effects of applying machine learning to more pages and by performing management operations at better granularity, compared to configurations of existing hybrid memory managers. Thaleia Dimitra Doudali, Ada Gavrilovska |
CCGRID | 2 |
| 2022 | Poster: Making Edge-assisted LiDAR Perceptions Robust to Lossy Point Cloud CompressionabstractReal-time light detection and ranging (LiDAR) perceptions, e.g., 3D object detection and simultaneous localization and mapping are computationally intensive to mobile devices of limited resources and often offloaded on the edge. Offloading Li-DAR perceptions requires compressing the raw sensor data, and lossy compression is used for efficiently reducing the data volume. Lossy compression degrades the quality of LiDAR point clouds, and the perception performance is decreased consequently. In this work, we present an interpolation algorithm improving the quality of a LiDAR point cloud to mitigate the perception performance loss due to lossy compression. The algorithm targets the range image (RI) representation of a point cloud and interpolates points at the RI based on depth gradients. Compared to existing image interpolation algorithms, our algorithm shows a better qualitative result when the point cloud is reconstructed from the interpolated RI. With the preliminary results, we also describe the next steps of the current work. Jin Heo, Gregoire Phillips, Per-Erik Brodin, Ada Gavrilovska |
SEC | 4 |
| 2022 | FLiCR: A Fast and Lightweight LiDAR Point Cloud Compression Based on Lossy RIabstractLight detection and ranging (LiDAR) sensors are becoming available on modern mobile devices and provide a 3D sensing capability. This new capability is beneficial for perceptions in various use cases, but it is challenging for resource-constrained mobile devices to use the perceptions in real-time because of their high computational complexity. In this context, edge computing can be used to enable LiDAR online perceptions, but offloading the perceptions on the edge server requires a low-latency, lightweight, and efficient compression due to the large volume of LiDAR point clouds data. This paper presents FLiCR, a fast and lightweight LiDAR point cloud compression method for enabling edge-assisted online perceptions. FLiCR is based on range images (RI) as an intermediate representation (IR), and dictionary coding for compressing RIs. FLiCR achieves its benefits by leveraging lossy RIs, and we show the efficiency of bytestream compression is largely improved with quantization and subsampling. In addition, we identify the limitation of current quality metrics for presenting the entropy of a point cloud, and introduce a new metric that reflects both point-wise and entropy-wise qualities for lossy IRs. The evaluation results show FLiCR is more suitable for edge-assisted real-time perceptions than the existing LiDAR compressions, and we demonstrate the effectiveness of our compression and metric with the evaluations on 3D object detection and LiDAR SLAM. Jin Heo, Christopher Phillips, Ada Gavrilovska |
SEC | 3 |
| 2022 | Poster: Fine-grained Control Plane Container Profiler for MECabstractToday, the edge computing system stack is built by leveraging the current cloud technologies, such as the containers, Kubernetes, etc., because, like the cloud, the edge is multi-tenant infrastructure. However, edge applications have more latency-critical SLAs and the infrastructure itself resource-constrained. That puts additional burdens on its control plane, which are not addressed by the cloud control plain tools. At the edge, if deployments aren't specified accurately, edge providers will face the dilemma between the waste of resource due to overcommitment vs. SLA violations. However, we observed that it is not feasible to rely on the existing monitoring tools, designed for the cloud, to glean that information from workloads with varying use of resources, at the needed fine granularity. Trying to do that with brute-forcing cloud solutions turns out to be extremely demanding on the resources allocated to the control plane. We present a new control plane tool, Colibri, aimed at addressing those conflicting requirements. Colibri can be dispatched dynamically, when needed, and enables characterization of containers deployed using Kubernetes across CPU, memory and network resource usage patterns at millisecond scale. The preliminary results demonstrate the effectiveness of out approach in reducing SLA violations by up to 98% for representative edge workloads. Ke-Jou Hsu, Ketan Bhardwaj, Ada Gavrilovska |
SEC | 3 |
| 2022 | ShapeShifter: Resolving the Hidden Latency Contention Problem in MECabstractMobile Edge Computing (MEC) creates new infrastructure at the edges of the mobile networks, thus providing transformative opportunities for applications seeking latency benefits by operating closer to end-users and devices. However, the reduced network distance between the application endpoints of the MEC flows causes pattern shifts in the packet bursts exchanged at the network edges. The longer and denser bursts create a new source of contention that is not considered by current solutions. As a result, naively collocating applications onto the MEC tier can negatively affect latency-critical workloads, resulting in up to 73% packets experiencing as much as 3.8x increased latency. This makes it impossible to support latency-centric SLOs in MEC, obviating its expected benefits from MEC. This paper is the first to describe this new contention point in mobile networks and its potentially crippling impact on the achievable latency benefit from MEC. We propose ShapeShifter, a new component in the MEC architecture which solves the MEC latency contention problem through adaptive latency-centric burst management of MEC flows. ShapeShifter is effective - it eliminates SLO violations for latency-critical applications and improves application performance in multi-tenant scenarios by up to 3.8 x – and practical – it can be deployed with minimal disruption to the current mobile network ecosystem. Valentin Rakovic, Ke-Jou Hsu, Ketan Bhardwaj, Ada Gavrilovska, Liljana Gavrilovska |
SEC | 4 |
| 2022 | FAM-Graph: Graph Analytics on Disaggregated MemoryabstractDisaggregated memory is being proposed as a way to provide efficient memory scaling for data intensive applications. High performance interconnect technologies, such as CXL, make disaggregated, fabric-attached-memory (FAM) a viable secondary tier of memory. Previous work on remote memory relies on extending kernel level paging to utilize FAM as an additional storage tier after local memory. These approaches have the advantage of exposing remote memory in application transparent ways that do not require code changes, but they incur large overheads due to the mismatch between the abstraction of a flat virtual address space and the reality of the tiered nature of FAM. In this paper, we present an alternative approach to remote memory based on application-specific objects. We design FAM-Graph - a semi-external graph processing system that leverages application-level properties, such as read only edge data, to efficiently tier data between local and remote memory, and prefetch remote data for local computation. Using several graph algorithms and datasets, we demonstrate that FAM-Graph achieves end-to-end performance within factors of 1–6× of Galois, the state of the art shared memory graph processing system, while using up to 20× less local memory. When Galois is used in conjunction with an OS-level FAM solution, we show that FAM-Graph achieves better end-to-end performance by up to 9× when both systems are configured with the same amount of local memory. Daniel Zahka, Ada Gavrilovska |
IPDPS | 2 |
| 2021 | Distributed Work Stealing at Scale via MatchmakingabstractMany classes of high-performance applications and combinatorial problems exhibit large degree of imbalance in terms of task execution times. One approach to achieving high machine efficiency and balanced resource use is to over-decompose the problem into fine-grained tasks that are then dynamically re-balanced across the system by approaches such as workstealing. For such irregular applications, existing work stealing or work balancing techniques exhibit high overheads due to potentially excessive communication messages, and/or delays experienced by idle nodes in finding work due to repeated failed steals. In particular, on very large scale clusters, the problem of matching work generation and work consumption is extremely challenging. At the heart of the problem is the dilemma: idle workers do not know where to look for work and the work producers do not know where to send the generated extra work. We contend that the fundamental problem of distributed work-stealing is not one of rapid dissemination of availability of work but one of rapidly bringing together work producers and consumers. In response, we develop a workstealing-based algorithm that performs timely, lightweight and highly efficient matchmaking between work producers and consumers. The matchmaker simultaneously accumulates the work availability messages from the producer nodes and rapidly redirects messages of work availability to the consumers. The matchmaker is a regular worker which works on its own work queue executing tasks when it pauses from the matchmaking work. We validate these claims via an implementation of a match-making-based scheduler for Charm++, that we evaluate using representative benchmarks running on up to 8K cores. Results show that our scheduler is able to outperform other load balancers and distributed work stealing schedulers, and to achieve scale beyond what is possible with current approaches. Hrushit Parikh, Vinit Deodhar, Ada Gavrilovska, Santosh Pande |
CLUSTER | 3 |
| 2021 | Machine Learning Augmented Hybrid Memory ManagementabstractThe integration of emerging non volatile memory hardware technologies into the main memory substrate, enables massive memory capacities at a reasonable cost in return for slower access speeds. This heterogeneity, along with the greater irregularity in the behavior of emerging workloads, render existing memory management approaches ineffective. This creates a significant gap between the realized vs. achievable performance and efficiency. At the same time, resource management solutions augmented with machine learning show great promise for fine-tuning system configuration knobs and predicting future behaviors. This thesis builds novel system-level mechanisms and reveals new insights towards the practical integration of machine learning in hybrid memory management. The specific contributions of this thesis is a machine learning augmented memory manager, coupled with insightful mechanisms to reduce the associated learning overheads and fine-tune critical operational parameters. The impact of this thesis is realizing an average of 3x application performance improvements and setting the new state-of-the-art in hybrid memory management. Thaleia Dimitra Doudali, Ada Gavrilovska |
HPDC | 2 |
| 2021 | The Performance Argument for Blockchain-based Edge DNS Caching
James Choncholas, Ketan Bhardwaj, Ada Gavrilovska |
SEC | 3 |
| 2021 | Poster: Enabling Flexible Edge-assisted XR
Jin Heo, Ketan Bhardwaj, Ada Gavrilovska |
SEC | 3 |
| 2021 | Cori: Dancing to the Right Beat of Periodic Data Movements over Hybrid Memory SystemsabstractEmerging hybrid memory systems that comprise technologies such as Intel's Optane DC Persistent Memory, exhibit disparities in the access speeds and capacity ratios of their heterogeneous memory components. This breaks many assumptions and heuristics designed for traditional DRAM-only platforms. High application performance is feasible via dynamic data movement across memory units, which maximizes the capacity use of DRAM while ensuring efficient use of the aggregate system resources. Newly proposed solutions use performance models and machine intelligence to optimize which and how much data to move dynamically. However, the decision of when to move this data is based on empirical selection of time intervals, or left to the applications. Our experimental evaluation shows that failure to properly conFigure the data movement frequency can lead to 10%-100% performance degradation for a given data movement policy; yet, there is no established methodology on how to properly conFigure this value for a given workload, platform and policy. We propose Cori, a system-level tuning solution that identifies and extracts the necessary application-level data reuse information, and guides the selection of data movement frequency to deliver gains in application performance and system resource efficiency. Experimental evaluation shows that Cori configures data movement frequencies that provide application performance within 3% of the optimal one, and that it can achieve this up to 5 x more quickly than random or brute-force approaches. System-level validation of Cori on a platform with DRAM and Intel's Optane DC PMEM confirms its practicality and tuning efficiency. Thaleia Dimitra Doudali, Daniel Zahka, Ada Gavrilovska |
IPDPS | 3 |
| 2021 | Introduction to the Special Issue on USENIX ATC 2020abstractNo abstract available. Ada Gavrilovska, Erez Zadok |
ACM Trans. Storage | 1 |
| 2020 | DNS Does Not Suffice for MEC-CDNabstractMobile edge computing (MEC) can transform mobile networks into a new infrastructure tier for services requiring low response times, such as those providing content to emerging AR/VR, autonomous driving, and other types of applications. To be successful, the CDNs operating in this MEC infrastructure tier MEC-CDNs will need to ensure end user applications gain access to a cache server in a fast and accurate manner. This paper sheds light on the challenges that the current mobile DNS architecture poses toward achieving this goal, and presents ideas on how to re-architect the existing DNS architecture to enable CDNs to provide low-latency content delivery from the edge. Ke-Jou Hsu, James Choncholas, Ketan Bhardwaj, Ada Gavrilovska |
HotNets | 4 |
| 2020 | Generating Robust Parallel Programs via Model Driven Prediction of Compiler Optimizations for Non-determinismabstractExecution orders in parallel programs are governed by non-determinism and can vary substantially across different executions even on the same input. Thus, a highly non-deterministic program can exhibit rare execution orders never observed during testing. It is desirable to reduce non-determinism to suppress corner case behavior in production cycle (making the execution robust or bug-free) and increase non-determinism for reproducing bugs in the development cycle. Performance-wise different optimization levels (e.g. from O0 to O3) are enabled during development , however, non-determinism-wise, developers have no way to select right compiler optimization level in order to increase non-determinism for debugging or to decrease it for robustness. Girish Mururu, Kaushik Ravichandran 0001, Ada Gavrilovska, Santosh Pande |
ICPP | 3 |
| 2019 | Quantifying and Reducing Execution Variance in STM via Model Driven Commit OptimizationabstractSimplified parallel programming coupled with an ability to express speculative computation is realized with Software Transactional Memory (STM). Although STMs are gaining popularity because of significant improvements in parallel performance, they exhibit enormous variation in transaction execution with non-repeatable performance behavior which is unacceptable in many application domains, especially in which frame rates and responsiveness should be predictable. In other domains reproducible transactional behavior helps towards system provisioning. Thus, reducing execution variance in STM is an important performance goal that has been mostly overlooked. In this work, we minimize the variance in execution time of threads in STM by reducing non-determinism exhibited due to speculation. We define the state of STM, and we use it to first quantity non-determinism and then generate an automaton that models the execution behavior of threads in STM. We finally use the state automaton to guide the STM to avoid non-predictable transactional behavior thus reducing non-determinism in roll-backs which in turn results in reduction in variance. We observed average reduction of variance in execution time of threads up to 74% in 16 cores and 53% in 8 cores by reducing non-determinism up to 24% in 16 cores and 44% in 8 cores, respectively, on STAMP benchmark suite while experiencing average slowdown of 4.8% in 8 cores and 19.2% in 16 cores. We also reduced the variance in frame rate by maximum of 65% on a version of real world game Quake3 without degradation in timing. Girish Mururu, Ada Gavrilovska, Santosh Pande |
CGO | 2 |
| 2019 | Cartel: A System for Collaborative Transfer Learning at the EdgeabstractAs Multi-access Edge Computing (MEC) and 5G technologies evolve, new applications are emerging with unprecedented capacity and real-time requirements. At the core of such applications there is a need for machine learning (ML) to create value from the data at the edge. Current ML systems transfer data from geo-distributed streams to a central datacenter for modeling. The model is then moved to the edge and used for inference or classification. These systems can be ineffective because they introduce significant demand for data movement and model transfer in the critical path of learning. Furthermore, a full model may not be needed at each edge location. An alternative is to train and update the models online at each edge with local data, in isolation from other edges. Still, this approach can worsen the accuracy of models due to reduced data availability, especially in the presence of local data shifts. Harshit Daga, Patrick K. Nicholson, Ada Gavrilovska, Diego Lugones |
SoCC | 3 |
| 2019 | Serving Mobile Apps: A Slice at a TimeabstractEnd users wanting to do more and more with mobile apps has led to explosive growth in the number of available apps. This has widened the gap between developers making apps available and end users being able to install all the apps they want on their device. To address this, Google introduced Instant Apps for Android where users can access selective app features on demand without having to download and install entire apps. But this requires developers to refactor apps and limits the apps' functionality. Ketan Bhardwaj, Matt Saunders, Nikita Juneja, Ada Gavrilovska |
EuroSys | 4 |
| 2019 | Kleio: A Hybrid Memory Page Scheduler with Machine IntelligenceabstractThe increasing demand of big data analytics for more main memory capacity in datacenters and exascale computing environments is driving the integration of heterogeneous memory technologies. The new technologies exhibit vastly greater differences in access latencies, bandwidth and capacity compared to the traditional NUMA systems. Leveraging this heterogeneity while also delivering application performance enhancements requires intelligent data placement. We present Kleio, a page scheduler with machine intelligence for applications that execute across hybrid memory components. Kleio is a hybrid page scheduler that combines existing, lightweight, history-based data tiering methods for hybrid memory, with novel intelligent placement decisions based on deep neural networks. We contribute new understanding toward the scope of benefits that can be achieved by using intelligent page scheduling in comparison to existing history-based approaches, and towards the choice of the deep learning algorithms and their parameters that are effective for this problem space. Kleio incorporates a new method for prioritizing pages that leads to highest performance boost, while limiting the resulting system resource overheads. Our performance evaluation indicates that Kleio reduces on average 80% of the performance gap between the existing solutions and an oracle with knowledge of future access pattern. Kleio provides hybrid memory systems with fast and effective neural network training and prediction accuracy levels, which bring significant application performance improvements with limited resource overheads, so as to lay the grounds for its practical integration in future systems. Thaleia Dimitra Doudali, Sergey Blagodurov, Abhinav Vishnu, Sudhanva Gurumurthi, Ada Gavrilovska |
HPDC | 5 |
| 2019 | Addressing the Fragmentation Problem in Distributed and Decentralized Edge Computing: A VisionabstractAt the core of the value proposition of edge computing is the ability to put computation close enough to the data sources, on demand. However, the data sources, computational infrastructure and software services needed to come together to power emerging and future edge computing applications are fragmented across different stakeholders, each with their own incentives, policies, and constraints on resources they can afford. This fragmentation limits the ability of edge computing to guarantee to applications and data the edge which will deliver the desired benefit. In this paper, we present our vision for an Edge Exchange, a decentralized directory service for a multi-stakeholder edge, as a path forward to enabling applications to be deployed across the best available edge resources, while still providing each stakeholder with controls regarding their resource use and sharing policies. Ketan Bhardwaj, Ada Gavrilovska, Vladimir Kolesnikov, Matt Saunders, Hobin Yoon, Mugdha Bondre, Meghana Babu, Jacob Walsh |
IC2E | 2 |
| 2019 | Collaboration Versus Cheating: Reducing Code Plagiarism in an Online MS Computer Science ProgramabstractWe outline how we detected programming plagiarism in an introductory online course for a master's of science in computer science program, how we achieved a statistically significant reduction in programming plagiarism by combining a clear explanation of university and class policy on academic honesty reinforced with a short but formal assessment, and how we evaluated plagiarism rates before and after implementing our policy and assessment. Tony Mason, Ada Gavrilovska, David A. Joyner |
SIGCSE | 2 |
| 2018 | Mnemo: Boosting Memory Cost Efficiency in Hybrid Memory SystemsabstractMnemo is an application profiling tool specialized for data serving and caching workloads, which retrieve data from cloud in-memory key-value stores. The increasing demand to boost application performance via in-memory data retrieval and the resulting spike in the overall system hosting cost, lead to the promise that cheaper but slower memory technologies, such as NVDIMMs (Non Volatile Memory), are going to co-exist with the currently predominant ones, i.e. DRAM. In such future cloud systems where the memory substrate is going to include heterogeneous hardware, Mnemo comes as the necessary memory sizing and data tiering consultant. Mnemo permits quick exploration of the trade-offs between the system cost and application performance, due to the various possible sizings of the hybrid memory system components. Thaleia Dimitra Doudali, Ada Gavrilovska |
SoCC | 2 |
| 2018 | Mutant: Balancing Storage Cost and Latency in LSM-Tree Data StoresabstractToday's cloud database systems are not designed for seamless cost-performance trade-offs for changing SLOs. Database engineers have a limited number of trade-offs due to the limited storage types offered by cloud vendors, and switching to a different storage type requires a time-consuming data migration to a new database. We propose Mutant, a new storage layer for log-structured merge tree (LSM-tree) data stores that dynamically balances database cost and performance by organizing SSTables (files that store a subset of records) into different storage types based on SSTable access frequencies. We implemented Mutant by extending RocksDB and found in our evaluation that Mutant delivers seamless cost-performance trade-offs with the YCSB workload and a real-world workload trace. Moreover, through additional optimizations, Mutant lowers the user-perceived latency significantly compared with the unmodified database. Hobin Yoon, Juncheng Yang, Sveinn Fannar Kristjansson, Steinn E. Sigurðarson, Ymir Vigfusson, Ada Gavrilovska |
SoCC | 6 |
| 2018 | NVStream: accelerating HPC workflows with NVRAM-based transport for streaming objectsabstractNonvolatile memory technologies (NVRAM) with larger capacity relative to DRAM and faster persistence relative to block-based storage technologies are expected to play a crucial role in accelerating I/O performance for HPC scientific workflows. Typically, a scientific workflow includes a simulation process (producer of data) and an analytics application process (consumer of data) that stream, share, and exchange data supported by an underlying OS-level file system. However, using an OS-level file system for data sharing adds substantial software overheads due to frequent system calls, journaling (for crash-consistency) cost, and file-system metadata update cost. To overcome these challenges, we design NVStream- a lightweight user-level data management system that exploits NVRAMs byte addressability and fast persistence to support streaming I/O in scientific workflows. First, NVStream reduces I/O-related software overheads by designing a memory-based persistent object store and log-structured heap manager that exploit NVRAM's large capacity. Second, NVStream incorporates a hardware-assisted non-temporal stores for crash-consistent updates at near hardware data copy (memory copy) speeds. Finally, NVStream reduces data written to NVRAM with a delta compression, which further reduces I/O cost for workflows with higher write locality. The evaluation of NVStream using I/O benchmarks and scientific applications demonstrates 10X reduction in I/O compared to NVRAM-optimized file systems and also guaranteeing crash-consistent data movement. Pradeep Fernando, Ada Gavrilovska, Sudarsun Kannan, Greg Eisenhauer |
HPDC | 2 |
| 2018 | Quantifying and reducing execution variance in STM via model driven commit optimizationabstractSimplified parallel programming coupled with an ability to express speculative computation is realized with Software Transactional Memory (STM). Although STMs are gaining popularity because of significant improvements in parallel performance, they exhibit enormous variation in transaction execution with non-repeatable performance behavior which is unacceptable in many application domains, especially in which frame rates and responsiveness should be predictable. Thus, reducing execution variance in STM is an important performance goal that has been mostly overlooked. In this work, we minimize the variance in execution time of threads in STM by reducing non-determinism exhibited due to speculation by first quantifying non-determinism and generating an automaton that models the behavior of STM. We used the automaton to guide the STM to a less non-deterministic execution that reduced the variance in frame rate by a maximum of 65% on a version of real-world Quake3 game. Girish Mururu, Ada Gavrilovska, Santosh Pande |
PPoPP | 2 |
| 2018 | Redesigning LSMs for Nonvolatile Memory with NoveLSM
Sudarsun Kannan, Nitish Bhat, Ada Gavrilovska, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
USENIX ATC | 3 |
| 2017 | Fault-Scalable Virtualized Infrastructure ManagementabstractLarge-scale virtualized datacenters require considerable automation in infrastructure management in order to operate efficiently. Automation is impaired, however, by the fact that deployments are prone to multiple types of subtle faults due to hardware failures, software bugs, misconfiguration, crashes, performance degraded hardware, etc. Existing Infrastructure-as-a-Service (IaaS) management stacks incorporate little to no resilience measures to shield end users from such cloud providerlevel failures and poor performance. This paper proposes and evaluates extensions to IaaS stacks that mask faults in a fault-agnostic manner while ensuring that the overheads can be proportional to observed failure rates. We also demonstrate that infrastructure automation services and end-user applications can use service-specific knowledge, together with our new interface, to achieve better outcomes. Mukil Kesavan, Ada Gavrilovska, Karsten Schwan |
ICDCS | 2 |
| 2017 | HeteroOS: OS Design for Heterogeneous Memory Management in Datacenter
Sudarsun Kannan, Ada Gavrilovska, Vishal Gupta 0001, Karsten Schwan |
ISCA | 2 |
| 2017 | Concurrent Log-Structured Memory for Many-Core Key-Value StoresabstractKey-value stores are an important tool in managing and accessing large in-memory data sets. As many applications benefit from having as much of their working state fit into main memory, an important design of the memory management of modern key-value stores is the use of log-structured approaches, enabling efficient use of the memory capacity, by compacting objects to avoid fragmented states. However, with the emergence of thousand-core and peta-byte memory platforms (DRAM or future storage-class memories) log-structured designs struggle to scale, preventing parallel applications from exploiting the full capabilities of the hardware: careful coordination is required for background activities (compacting and organizing memory) to remain asynchronous with respect to the use of the interface, and for insertion operations to avoid contending for centralized resources such as the log head and memory pools. In this work, we present the design of a log-structured key-value store called Nibble that incorporates a multi-head log for supporting concurrent writes, a novel distributed epoch mechanism for scalable memory reclamation, and an optimistic concurrency index. We implement Nibble in the Rust language in ca. 4000 lines of code, and evaluate it across a variety of data-serving workloads on a 240-core cache-coherent server. Our measurements show Nibble scales linearly in uniform YCSB workloads, matching competitive non-log-structured key-value stores for write- dominated traces at 50 million operations per second on 1 TiB-sized working sets. Our memory analysis shows Nibble is efficient, requiring less than 10% additional capacity, whereas memory use by non-log-structured key-value store designs may be as high as 2x. Alex Merritt, Ada Gavrilovska, Yuan Chen 0001, Dejan S. Milojicic |
Proc. VLDB Endow. | 2 |
| 2016 | Energy Aware Persistence: Reducing Energy Overheads of Memory-based Persistence in NVMsabstractNext generation byte addressable nonvolatile memories (NVMs) such as PCM, Memristor, and 3D X-Point are attractive solutions for mobile and other end-user devices, as they offer memory scalability as well as fast persistent storage. However, NVM's limitations of slow writes and high write energy are magnified for applications that require atomic, consistent, isolated and durable (ACID) persistence. For maintaining ACID persistence guarantees, applications not only need to do extra writes to NVM but also need to execute a significant number of additional CPU instructions for performing NVM writes in a transactional manner. Our analysis shows that maintaining persistence with ACID guarantees increases CPU energy up to 7.3x and NVM energy up to 5.1x compared to a baseline with no ACID guarantees. For computing platforms such as mobile devices, where energy consumption is a critical factor, it is important that the energy cost of persistence is reduced. Sudarsun Kannan, Moinuddin K. Qureshi, Ada Gavrilovska, Karsten Schwan |
PACT | 3 |
| 2016 | pVM: persistent virtual memory for efficient capacity scaling and object storageabstractNext-generation byte-addressable nonvolatile memories (NVMs), such as phase change memory (PCM) and Memristors, promise fast data storage, and more importantly, address DRAM scalability issues. State-of-the-art OS mechanisms for NVMs have focused on improving the block-based virtual file system (VFS) to manage both persistence and the memory capacity scaling needs of applications. However, using the VFS for capacity scaling has several limitations, such as the lack of automatic memory capacity scaling across DRAM and NVM, inefficient use of the processor cache and TLB, and high page access costs. These limitations reduce application performance and also impact applications that use NVM for persistent object storage with flat namespaces, such as photo stores, NoSQL databases, and others. Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan |
EuroSys | 2 |
| 2016 | Phoenix: Memory Speed HPC I/O with NVMabstractIn order to bridge the gap between the applications' I/O needs on future exascale platforms, and thecapabilities of conventional memory and storage technologies, HPC system designs started integrating components based onemerging non-volatile memory technologies. Non-volatile memory (NVRAM) provides persistent storage at close to memoryspeeds, with good capacity scaling, leading to opportunitiesto accelerate I/O in exascale machines. However, naive use ofNVRAM devices with current software stacks, exposes newbottlenecks due to the limited device bandwidth and slowerdevice access times compared to DRAM. To address this, we propose Phoenix (PHX), an NVRAM-bandwidth aware object store for persistent objects. PHXachieves efficiency through use of memory-centric objectinterfaces and device access stack specialized for NVRAM. Furthermore, PHX deals with the limited PCM bandwidththrough simultaneous use of NVRAM and DRAM devices, thus increasing the effective data movement bandwidth. Thisleads to reduction in the time length of the critical path I/Ooperations associated with the slow NVM device. To continueguaranteeing adequate reliability for the persistent objects, DRAM-resident object state is replicated across peer nodes'memory, accessible through high-bandwidth interconnects. Furthermore PHX minimizes the data movement overheads dueto additional data copies, by using a cost model that considersdevice bandwidths, remote storage distance and energy costs. Experimental analysis using real-world HPC applications onemulated NVRAM hardware shows that Phoenix's controlleduse of node-local and remote-node memory bandwidth, delivers up to ~ 1.2×, ~ 2× and ~ 12× speed-up for checkpoint I/Ofor the S3D, CM1 and GTC HPC applications, respectively. Furthermore PHX reduces total simulation checkpoint over-head of GTC up to ~ 18%. Pradeep Fernando, Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan |
HiPC | 3 |
| 2016 | Implications of Heterogeneous Memories in Next Generation Server SystemsabstractNext generation datacenter and exascale machines will include significantly larger amounts of memory, greater heterogeneity in the performance, persistence or sharing properties of the memory components they encompass, and increase in the relative cost and complexity of the data paths in the resulting memory topology. This poses several challenges to the systems software stacks managing these memory-centric platform designs. First, technology advances in novel memory technologies shift the data access bottlenecks into the software stack. Second, current systems software lacks capabilities to bridge the multi-dimensional non-uniformity in the memory subsystem to the dynamic nature of the workloads it must support. In addition, current memory management solutions have limited ability to explicitly reason about the costs and tradeoffs associated with data movement operations, leading to limited efficiency of their interconnect use. To address these problems, next generation systems software stacks require new data structures, abstractions and mechanisms in order to enable new levels of efficiency in the data placement, movement, and transformation decisions that govern the underlying memory use. In this talk, I will present our approach to rearchitecting systems software and services in response to both node-level and system-wide memory heterogeneity and scale, particularly concerning the presence of non-volatile memories, and will demonstrate the resulting performance and efficiency gains using several scientific and data-intensive workloads. Ada Gavrilovska |
HPDC | 1 |
| 2016 | Attribute-Based Partial Geo-Replication SystemabstractExisting partial geo-replication systems do not always provide optimal cost or latency, because their replication decisions are based on statically established data access popularity metrics, regardless of the application types. We demonstrate that additional reduction in cost and latency can be achieved by (1) using the right object attributes for making replication decisions for each type of application, (2) using multi-attribute-based replications, and (3) combining the popularity-based but reactive approach with the more random but proactive approach to data replication. Toward this end, we propose Acorn, an Attribute-based COntinuous partial geo-ReplicatioN system, and its prototype implementation based on Apache Cassandra. Experiments with two types of global-scale, data-sharing applications demonstrate up to 54% and 90% cost overhead reduction over existing systems or 38% and 91% latency overhead reduction. Hobin Yoon, Ada Gavrilovska, Karsten Schwan |
IC2E | 2 |
| 2016 | TCP Ordo: The cost of ordered processing in TCP serversabstractTo achieve scalable, high-throughput, low-latency packet processing, TCP implementations are parallelized across cores in multicore platforms. This, however, significantly affects the order in which packets from different flows are delivered to application processing. Our measurements record up to 75% of the packets are delivered to applications in a way that does not match the order in which they are received on the network interface. For many important classes of applications, such as financial services, bidding and trading engines, game engines, this cross-flow packet reordering affects their ability to provide fairness guarantees. To address this gap, we propose TCP-Ordo - a TCP stack which provides strict ordering of packets across multiple flows, as well as flexibility to control the degree to which ordering is enforced. TCP-Ordo outperforms existing TCP implementations in both latency and throughput. The current prototype is implemented as a user-level TCP stack for Mellanox Connectx3 NICs with 40Gbps Ethernet interfaces. Without ordering guarantees TCP-Ordo delivers one way latency of 4.75usec (a 5x improvement over the Linux kernel) for 150B packets and throughput of 22Gbps (a 2x improvement over mTCP) for 1500B packets. Furthermore, with TCP-Ordo applications are provided with strict packet order delivery, with performance that continues to be superior to the state-of-the-art, even when enforcing packet order across 800,000 connections on a 12-core platform. Finally, TCP-Ordo supports the notion of `ordered domains' that offer flexibility in the degree of ordering that an application will experience, and pay for. Mohan Kumar, Ada Gavrilovska |
INFOCOM | 2 |
| 2016 | Efficient distributed workstealing via matchmakingabstractMany classes of high-performance applications and combinatorial problems exhibit large degree of runtime load variability. One approach to achieving balanced resource use is to over decompose the problem on fine-grained tasks that are then dynamically balanced using approaches such as workstealing. Existing work stealing techniques for such irregular applications, running on large clusters, exhibit high overheads due to potential untimely interruption of busy nodes, excessive communication messages and delays experienced by idle nodes in finding work due to repeated failed steals. We contend that the fundamental problem of distributed work-stealing is of rapidly bringing together work producers and consumers. In response, we develop an algorithm that performs timely, lightweight and highly efficient matchmaking between work producers and consumers which results in accurate load balance. Experimental evaluations show that our scheduler is able to outperform other distributed work stealing schedulers, and to achieve scale beyond what is possible with current approaches. Hrushit Parikh, Vinit Deodhar, Ada Gavrilovska, Santosh Pande |
PPoPP | 3 |
| 2016 | Virtualizing the Edge of the Cloud: the New FrontierabstractOver the last two decades, virtualization technologies have turned datacenter infrastructure into multitenant, dynami- cally provisionable, elastic resource, and formed the basis for the wide adoption of cloud computing. Many of todays cloud applications, however, are based on continuous inter- actions with end users and their devices, and the trend is only expected to intensify with the expansion of the Internet of Things. The consequent bandwidth and latency require- ments of these emerging workloads push the cloud bound- ary outside of traditional datacenters, giving rise to an edge tier in the end-device-to-cloud-backend infrastructure. Com- putational resources embedded in anything from standalone microservers to WiFi routers and small cell access points, and their open APIs, present opportunities for deploying ap- plication logic and state closer to where it is being used, addressing both latency and backhaul bandwidth problems. This talk will look at the role that existing virtualization tech- nologies can play in providing in this edge tier the required flexibility, dynamic provisioning and isolation, and will out- line open problems that require development of new solu- tions. We will also discuss the opportunities to leverage these technologies to further deal with the diversity in the end-user device and IoT space. Ada Gavrilovska |
VEE | 1 |
| 2015 | Compiler Assisted Load Balancing on Large ClustersabstractLoad balancing of tasks across processing nodes is critical for achieving speed up on large scale clusters. Load balancing schemes typically detect the imbalance and then migrate the load from an overloaded processing node to an idle or lightly loaded processing node and thus, estimation of load critically affects the performance of load balancing schemes. On large scale clusters, the latency of load migration between processing nodes (and the energy) is also a significant overhead and any missteps in load estimation can cause significant migrations and performance losses. Currently, the load estimation is done either by profile or feedback driven approaches in sophisticated systems such as Charm++, but such approaches must be re-thought in light of some workloads such as Adaptive Mesh Refinement (AMR) and multiscale physics where the load variations could be quite dynamic and rapid. In this work we propose a compiler based framework which performs precise prediction of the forthcoming workload. The compiler driven load prediction technique performs static analysis of a task and derives an expression to predict load of a task and hoists it as early as possible in the control flow of execution. The compiler also inserts corrector expressions at strategic program points which refine the reachability probability of the load as well as its estimation. The predictor and the corrector expressions are evaluated at runtime and the predicted load information is refined as the execution proceeds and is eventually used by load balancer to take efficient migration decisions. We present an implementation of the above in the Rose compiler and the Charm++ parallel programming framework. We demonstrate the effectiveness of the framework on some key benchmarks that exhibit dynamic variations and show how the compiler framework assists load balancing schemes in Charm++ to provide significant gains. Vinit Deodhar, Hrushit Parikh, Ada Gavrilovska, Santosh Pande |
PACT | 3 |
| 2014 | DeSTM: harnessing determinism in STMs for application developmentabstractNon-determinism has long been recognized as one of the key challenges which restrict parallel programmer productivity by complicating several phases of application development. While Software Transactional Memory (STM) systems have greatly improved the productivity of programmers developing parallel applications in a host of areas they still exhibit non-deterministic behavior leading to decreased productivity. While determinism in parallel applications which use traditional synchronization primitives (such as locks) has been relatively well studied, its interplay with STMs has not. In this paper we present DeSTM, a deterministic STM, which allows programmers to leverage determinism through the implementation, debugging and testing phases of application development. Kaushik Ravichandran 0001, Ada Gavrilovska, Santosh Pande |
PACT | 2 |
| 2014 | Merlin: Application- and Platform-aware Resource Allocation in Consolidated Server SystemsabstractWorkload consolidation, whether via use of virtualization or with lightweight, container-based methods, is critically important for current and future datacenter and cloud computing systems. Yet such consolidation challenges the ability of current systems to meet application resource needs and isolate their resource shares, particularly for high core count or 'scaleup' servers. This paper presents the 'Merlin' approach to managing the resources of multicore platforms, which satisfies an application's resource requirements efficiently -- using low cost allocations -- and improves isolation -- measured as increased predictability of application execution. Merlin (i) creates a virtual platform (VP) as a system-level resource commitment to an application's resource shares, (ii) enforces its isolation, and (iii) operates with low runtime overhead. Further, Merlin's resource (re)-allocation and isolation methods operate by constructing online models that capture the resource 'sensitivities' of the currently running applications along all of their resource dimensions. Elevating isolation into a first-class management principle, these sensitivity- and cost-based allocation and sharing methods lead to efficient methods for shared resource use on scaleup server systems. Experimental evaluations on a large core-count machine demonstrate improved performance with reduced performance variation and increased system throughput and efficiency, for a wide range of popular datacenter workloads, compared with the methods used in prior work and with the state-of-art Xen hypervisor. Priyanka Tembey, Ada Gavrilovska, Karsten Schwan |
SoCC | 2 |
| 2014 | HeteroCheckpoint: Efficient Checkpointing for Accelerator-Based SystemsabstractMoving toward exascale, the number of GPUs in HPC machines is bound to increase, and applications will spend increasing amounts of time running on those GPU devices. While GPU usage has already led to substantial speedup for HPC codes, their failure rates due to overheating are at least 10 times higher than those seen for the CPUs now commonly used on HPC machines. This makes it increasingly important for GPUs to have robust checkpoint/restart mechanisms. This paper introduces a unified CPU-GPU checkpoint mechanism, which can efficiently checkpoint the combined GPU-CPU memory state resident on machine nodes. Efficiency is gained in part by addressing the end-to-end data movements required for check pointing - from GPU to storage - by introducing novel pre-copy and checksum methods. These methods reduce checkpoint data movement cost seen by HPC applications, with initial measurements using different benchmark applications showing up to 60% reduced checkpoint overhead. Additional exploration of the use of next-generation storage, like NVM, show further promises of reduced check pointing overheads. Sudarsun Kannan, Naila Farooqui, Ada Gavrilovska, Karsten Schwan |
DSN | 3 |
| 2014 | Balancing context switch penalty and response time with elastic time slicingabstractVirtualization allows the platform to have increased number of logical processors by multiplexing the underlying resources across different virtual machines. The hardware resources get time shared not only between different virtual machines, but also between different workloads of the same virtual machine. An important source of performance degradation in such a scenario comes from the cache warmup penalties a workload experiences when it gets scheduled, as the working set belonging to the workload gets displaced by other concurrently running workloads. We show that a virtual machine that time switches between four workloads can cause some of the workloads a slowdown of as much as 54%. However, such performance degradation depends on the workload behavior, with some workloads experiencing negligible degradation and some severe degradation. We propose Elastic Time Slicing (ETS) to reduce the context switch overhead for the most affected workloads. We demonstrate that by taking the workload-specific context switch overhead into consideration, the CPU scheduler can make better decisions to minimize the context switch penalty for the most affected workloads, thereby resulting in substantial performance improvements. ETS enhances performance without compromising on response time, thereby achieving dual benefits. To facilitate ETS, we develop a low-overhead hardware-based mechanism that dynamically estimates the sensitivity of a given workload to context switching. We evaluate the accuracy of the mechanism under various cache management policies and show that it is very reliable. Context switch related warmup penalties increase as optimizations are applied to address traditional cache misses. For the first time, we assess the impact of advanced replacement policies and establish that it is significant. Nagakishore Jammula, Moinuddin K. Qureshi, Ada Gavrilovska, Jongman Kim |
HiPC | 3 |
| 2014 | Reducing the cost of persistence for nonvolatile heaps in end user devicesabstractThis paper explores the performance implications of using future byte addressable non-volatile memory (NVM) like PCM in end client devices. We explore how to obtain dual benefits - increased capacity and faster persistence - with low overhead and cost. Specifically, while increasing memory capacity can be gained by treating NVM as virtual memory, its use of persistent data storage incurs high consistency (frequent cache flushes) and durability (logging for failure) overheads, referred to as `persistence cost'. These not only affect the applications causing them, but also other applications relying on the same cache and/or memory hierarchy. This paper analyzes and quantifies in detail the performance overheads of persistence, which include (1) the aforementioned cache interference as well as (2) memory allocator overheads, and finally, (3) durability costs due to logging. Novel solutions to overcome such overheads include (1) a page contiguity algorithm that reduces interference-related cache misses, (2) a cache efficient NVM write aware memory allocator that reduces cache line flushes of allocator state by 8X, and (3) hybrid logging that reduces durability overheads substantially. With these solutions, experimental evaluations with different end user applications and SPEC2006 benchmarks show up to 12% reductions in cache misses, thereby reducing the total number of NVM writes. Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan |
HPCA | 2 |
| 2014 | Personal clouds: Sharing and integrating networked resources to enhance end user experiencesabstractEnd user experiences on mobile devices with their rich sets of sensors are constrained by limited device battery lives and restricted form factors, as well as by the `scope' of the data available locally. The `Personal Cloud' distributed software abstractions address these issues by enhancing the capabilities of a mobile device via seamless use of both nearby and remote cloud resources. In contrast to vendor-specific, middleware-based cloud solutions, Personal Cloud instances are created at hypervisor-level, to create for each end user the federation of networked resources best suited for the current environment and use. Specifically, the Cirrostratus extensions of the Xen hypervisor can federate a user's networked resources to establish a personal execution environment, governed by policies that go beyond evaluating network connectivity to also consider device ownership and access rights, the latter managed in a secure fashion via standard Social Network Services. Experimental evaluations with both Linux- and Android-based devices, and using Facebook as the SNS, show the approach capable of substantially augmenting a device's innate capabilities, improving application performance and the effective functionality seen by end users. Minsung Jang, Karsten Schwan, Ketan Bhardwaj, Ada Gavrilovska, Adhyas Avasthi |
INFOCOM | 4 |
| 2013 | Distributed resource exchange: Virtualized resource management for SR-IOV InfiniBand clustersabstractThe commoditization of high performance interconnects, like 40+ Gbps InfiniBand, and the emergence of low-overhead I/O virtualization solutions based on SR-IOV, is enabling the proliferation of such fabrics in virtualized datacenters and cloud computing platforms. As a result, such platforms are better equipped to execute workloads with diverse I/O requirements, ranging from throughput-intensive applications, such as `big data' analytics, to latency-sensitive applications, such as online applications with strict response-time guarantees. Improvements are also seen for the virtualization infrastructures used in data center settings, where high virtualized I/O performance supported by high-end fabrics enables more applications to be configured and deployed in multiple VMs - VM ensembles (VMEs) - distributed and communicating across multiple datacenter nodes. A challenge for I/O-intensive VM ensembles is the efficient management of the virtualized I/O and compute resources they share with other consolidated applications, particularly in lieu of VME-level SLA requirements like those pertaining to low or predictable end-to-end latencies for applications comprised of sets of interacting services. This paper addresses this challenge by presenting a management solution able to consider such SLA requirements, by supporting diverse SLA-aware policies, such as those maintaining bounded SLA guarantees for all VMEs, or those that minimize the impact of misbehaving VMEs. The management solution, termed Distributed Resource Exchange (DRX), borrows techniques from principles of microeconomics, and uses online resource pricing methods to provide mechanisms for such distributed and coordinated resource management. DRX and its mechanisms allow policies to be deployed on such a cluster in order to provide SLA guarantees to some applications by charging all the interfering VMEs `equally' or based on the `hurt', i.e. amount of I/O performed by the VMEs. While these mechanisms are general, our implementation is specifically for SR-IOV-based fabrics like InfiniBand and the KVM hypervisor. Our experimental evaluation consists of workloads representative of data-analytics, transactional and parallel benchmarks. The results demonstrate the feasibility of DRX and its utility to maintain SLA for transactional applications. We also show that the impact to the interfering workloads is also within acceptable bounds for certain policies. Adit Ranadive, Ada Gavrilovska, Karsten Schwan |
CLUSTER | 2 |
| 2013 | Optimizing Checkpoints Using NVM as Virtual MemoryabstractRapid checkpointing will remain key functionality for next generation high end machines. This paper explores the use of node-local nonvolatile memories (NVM) such as phase-change memory, to provide frequent, low overhead checkpoints. By adapting existing multi-level checkpoint techniques, we devise new methods, termed NVM-checkpoints, that efficiently store checkpoints on both local and remote node NVM. The checkpoint frequencies are guided by failure models that capture the expected accessibility of such data after failure. To lower overheads, NVM-checkpoints reduce the NVM and interconnect bandwidth used with a novel pre-copy mechanism, which incrementally moves checkpoint data from DRAM to NVM before a local checkpoint is started. This reduces local checkpoint cost by limiting the instantaneous data volume moved at checkpoint time, thereby freeing bandwidth for use by applications. In fact, the pre-copy method can reduce peak interconnect usage up to 46%. Since our approach treats NVM as memory rather than as 'Ramdisk', pre-copying can be generalized to directly move data to remote NVMs. This results in 40% faster application execution times compared to asynchronous approaches not using pre-copying. Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan, Dejan S. Milojicic |
IPDPS | 2 |
| 2013 | Practical Compute Capacity Management for Virtualized DatacentersabstractWe present CCM (Cloud Capacity Manager) - a prototype system and its methods for dynamically multiplexing the compute capacity of virtualized datacenters at scales of thousands of machines, for diverse workloads with variable demands. Extending prior studies primarily concerned with accurate capacity allocation and ensuring acceptable application performance, CCM also sheds light on the tradeoffs due to two unavoidable issues in large scale commodity datacenters: (i) maintaining low operational overhead given variable cost of performing management operations necessary to allocate resources, and (ii) coping with the increased incidences of these operations' failures. CCM is implemented in an industry-strength cloud infrastructure built on top of the VMware vSphere virtualization platform and is currently deployed in a 700 physical host datacenter. Its experimental evaluation uses production workload traces and a suite of representative cloud applications to generate dynamic scenarios. Results indicate that the pragmatic cloud-wide nature of CCM provides up to 25% more resources for workloads and improves datacenter utilization by up to 20%, compared to the common alternative approach of multiplexing capacity within multiple independent smaller datacenter partitions. Mukil Kesavan, Irfan Ahmad 0005, Orran Krieger, Ravi Soundararajan, Ada Gavrilovska, Karsten Schwan |
IEEE Trans. Cloud Comput. | 5 |
| 2012 | Interactive Use of Cloud Services: Amazon SQS and S3abstractInteractive use of cloud services is of keen interest to science end users, including for storing and accessing shared data sets. This paper evaluates the viability of interactively using two important cloud services offered by Amazon: SQS (Simple Queue Service) and S3 (Simple Storage Service). Specifically, we first measure the send-to-receive message latencies of SQS and then determine and devise rate controls to obtain suitable latencies and latency variations. Second, for S3, when transferring data into the cloud, we determine that increased parallelism in Transfer Manager can significantly improve upload performance, achieving up to 4 times improvements with careful elimination of upload bottlenecks. Hobin Yoon, Ada Gavrilovska, Karsten Schwan, Jim Donahue |
CCGRID | 2 |
| 2011 | HEaRS: A Hierarchical Energy-Aware Resource Scheduler for Virtualized Data CentersabstractWith the increasing popularity of Internet-based cloud services, energy efficiency in large-scale Internet data centers has become important not only to curtail energy costs and alleviate environmental concern, but also because such systems can quickly reach the limits of power available to them. This paper investigates to what extent and how energy usage improvements through consolidation can benefit from taking into account the environmental influences and effects seen in data center systems. Toward that end, we present experimental results obtained in a fully instrumented, small scale data center and then use these results to propose a hierarchical energy-aware resource scheduler (HEaRS) for cluster workload placement and server provisioning, also considers the physical environment in which data center systems operate. Specifically, at the rack level, HEaRS tries to maintain a 'thermal balance' across the rack to avoid hot spots and reduce cooling costs. At the chassis level, HEaRS utilizes the proportional plus integral controller to achieve a balance in the levels of usage of electrical current between the two power domains in the chassis, which helps the chassis reach its most energy efficient state. Finally, at server level, HEaRS can employ known methods like dynamic voltage and frequency scaling or core idling to reduce power consumption. This results in a hierarchical set of controllers that jointly, implement holistic solutions to energy-aware resource scheduling for an entire rack, and this hierarchical solution can then be further extended to entire data centers. Our initial experiment result show opportunities for gains, with up to 16% in energy usage compared to methods that are not aware of the physical environment and up to 15% improvements in application performance. Meina Song, Junde Song, Ada Gavrilovska, Karsten Schwan |
CLUSTER | 4 |
| 2011 | ResourceExchange: Latency-Aware Scheduling in Virtualized Environments with High Performance FabricsabstractVirtualized infrastructures have seen strong acceptance in data center systems and applications, but have not yet seen adoptance for latency-sensitive codes which require I/O to arrive predictability, or response times to be generated within certain timeliness guarantees. Examples of such applications include certain classes of parallel HPC codes, server systems performing phonecall or multimedia delivery, or financial services in electronic trading platforms, like ICE and CME. In this paper, we argue that the use of high-performance, VMM-bypass capable devices can help create the virtualized infrastructures needed for the latency-sensitive applications listed above. However, to enable consolidation, problems to be solved go beyond efficient I/O virtualization, and include dealing with the shared use of I/O and compute resource, in ways that minimize or eliminate interference. Toward this end, we describe ResEx -- a resource management approach for virtualized RDMA-based platforms which incorporates concepts from supply-demand theory and congestion pricing to dynamically control the allocation of CPU and I/O resources of guest VMs. ResEx and its mechanisms and abstractions allow multiple 'pricing policies' to be deployed on these types of virtualized platforms, including such which reduce interference and enhance isolation by identifying and taxing VMs responsible for resource congestion. While the main ideas behind ResEx are more general, the design presented in this paper is specific for InfiniBand RDMA-based virtualized platforms due to the use of asynchronous monitoring needed to determine the VMs' I/O usage, and the methods to establish the trading rate for the underlying CPU and I/O resources. The latter is particularly necessary since the hypervisor's only mechanism to control I/O usage is by making appropriate adjustments in the VM's CPU resources. The experimental evaluation of our solution uses InfiniBand platforms virtualized with the open source Xen hyper visor, and an RDMA-based latency-sensitive benchmark, BenchEx, based on a model of a financial trading platform. The results demonstrate the utility of the ResEx approach in making RDMA-based virtualized platforms more manageable and better suited for hosting even latency-sensitive workloads. ResEx can reduce the latency interference by as much as 30% in some cases as shown. Adit Ranadive, Ada Gavrilovska, Karsten Schwan |
CLUSTER | 2 |
| 2011 | Cloud4Home - Enhancing Data Services with @Home CloudsabstractMobile devices, net books and laptops, and powerful home PCs are creating ever-growing computational capacity at the periphery of the Internet, and this capacity is supporting an increasingly rich set of services, including media-rich entertainment and social networks, gaming, home security applications, flexible data access and storage, and others. Such 'at the edge' capacity raises the question, however, about how to combine it with the capabilities present in the cloud computing infrastructures residing in data center systems and reachable via the Internet. The Cloud4Home project and approach presented in this paper addresses this topic, by enabling and exploring the aggregate use of @home and @datacenter computational and storage capabilities. Cloud4Home uses virtualization technologies to create content storage, access, and sharing services that are fungible both in terms of where stored objects are located and in terms of where they are manipulated. In this fashion, data services can provide low latency response to @home events as well as high throughput response when the higher and less predictable latencies of data center access can be tolerated. Cloud4Home is implemented with the Xen open source hypervisors for standard x86-based mobile to server platforms, and is evaluated using sample applications based on home security and video conversion services. Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan |
ICDCS | 2 |
| 2010 | FaReS: Fair Resource Scheduling for VMM-Bypass InfiniBand DevicesabstractIn order to address the high performance I/O needs of HPC and enterprise applications, modern interconnection fabrics, such as InfiniBand and more recently, 10GigE, rely on network adapters with RDMA capabilities. In virtualized environments, these types of adapters are configured in a manner that bypasses the hypervisor and allows virtual machines (VMs) direct device access, so that they deliver near-native low-latency/high-bandwidth I/O. One challenge with the bypass approach is that it causes the hypervisor to lose control over VM-device interactions, including the ability to monitor such interactions and to ensure fair resource usage by VMs. Fairness violations, however, permit low-priority VMs to affect the I/O allocations of other higher priority VMs and more geerally, lack of supervision can lead to inefficiencies in the usage of platform resources. This paper describes the FaReS system-level mechanisms for monitoring VMs' usage of bypass I/O devices. Monitoring information acquired with FaReS is then used to adjust VMM-level scheduling in order to improve resource utilization and/or ensure fairness properties across the sets of VMs sharing platform resources. FaReS employs a memory introspection-based tool for asynchronously monitoring VMM-bypass devices, using InfiniBand HCAs as a concrete example. FaReS and its very low overhead (<;1%) monitoring methods are evaluated with microbenchmarks and representative HPC codes running on modern multicore platforms connected via InfiniBand, using the Xen hypervisor. For these codes fairness is achieved within 2% of the required value. Adit Ranadive, Ada Gavrilovska, Karsten Schwan |
CCGRID | 2 |
| 2010 | Differential virtual time (DVT): rethinking I/O service differentiation for virtual machinesabstractThis paper investigates what it entails to provide I/O service differentiation and performance isolation for virtual machines on individual multicore nodes in cloud platforms. Sharing I/O between VMs is fundamentally different from sharing I/O between processes because guest VM operating systems use adaptive resource management mechanisms like TCP congestion avoidance, disk I/O schedulers, etc. The problem is that these mechanisms are generally sensitive to the magnitude and rate of change of service latencies, where failing to address these latency concerns while designing a service differentiation framework for I/O results in undue performance degradation and hence, insufficient isolation between VMs. This problem is addressed by the notion of Differential Virtual Time (DVT), which can provide service differentiation with performance isolation for VM guest OS resource management mechanisms. DVT is realized within a proportional share I/O scheduling framework for the Xen hypervisor, and its use requires no changes to guest OSs. DVT is applied to message-based I/O, but is also applicable to subsystems like disk I/O. Experimental results with DVT-based I/O scheduling for representative applications demonstrate the utility and effectiveness of the approach. Mukil Kesavan, Ada Gavrilovska, Karsten Schwan |
SoCC | 2 |
| 2008 | Active CoordinaTion (ACT) - toward effectively managing virtualized multicore cloudsabstractA key benefit of utility data centers and cloud computing infrastructure is the level of consolidation they can offer to arbitrary guest applications, and the substantial saving in operational costs and resources that can be derived in the process. However, significant challenges remain before it becomes possible to effectively and at low cost manage virtualized systems, particularly in the face of increasing complexity of individual many-core platforms, and given the dynamic behaviors and resource requirements exhibited by cloud guest VMs. This paper describes the active coordination (ACT) approach, aimed to address a specific issue in the management domain, which is the fact that management actions must (1) typically touch upon multiple resources in order to be effective, and (2) must be continuously refined in order to deal with the dynamism in the platform resource loads. ACT relies on the notion of class-of-service, associated with (sets of) guest VMs, based on which it maps VMs onto platform units, the latter encapsulating sets of platform resources of different types. Using these abstractions, ACT can perform active management in multiple ways, including a VM-specific approach and a black box approach that relies on continuous monitoring of the guest VMs' runtime behavior and on an adaptive resource allocation algorithm, termed Multiplicative Increase, Subtractive Decrease Algorithm with Wiggle Room. In addition, ACT permits explicit external events to trigger VM or application-specific resource allocations, e.g., leveraging emerging standards such as WSDM. The experimental analysis of the ACT prototype, built for Xen-based platforms, use industry-standard benchmarks, including RUBiS, Hadoop, and SPEC. They demonstrate ACT's ability to efficiently manage the aggregate platform resources according to the guest VMs' relative importance (class-of-service), for both the black-box and the VM-specific approach. Mukil Kesavan, Adit Ranadive, Ada Gavrilovska, Karsten Schwan |
CLUSTER | 3 |
| 2008 | ShareStreams-V: A Virtualized QoS Packet Scheduling AcceleratorabstractThis paper introduces a virtualized FPGA-based accelerator for wire speed scheduling of packet streams under quality of service constraints. This work implements the dynamic window constrained scheduling algorithm and builds upon our previous custom accelerator by adding support for virtualization. This implementation is parametric, permitting tradeoffs between packet decision latency, decision throughput, and the number of virtual packet schedulers supported. When scheduling streams from multiple processes, ShareStreams-V 1 is able to schedule minimal size packets faster than one decision per 51.2 ns for up to 64 streams, the throughput required for 10 gbps Ethernet. The bottleneck currently is the host-accelerator HW/SW (PCIe) interface; this may be mitigated using high-speed interconnects/interfaces such as HyperTransport. Kangtao Kendall Chuang, Sudhakar Yalamanchili, Ada Gavrilovska, Karsten Schwan |
FCCM | 3 |
| 2008 | Flexible Classification on Heterogenous Multicore Appliance PlatformsabstractEmerging heterogeneous multicore systems are suitable platforms for efficient deployment of application- specific service components. Easily virtualizable and reprogrammable platforms, these 'appliances' for future information services make it possible to run familiar software stacks created with common development tools, in addition to offering acceleration capabilities for performance critical software components. Of particular importance are the efficient and scalable execution of the content-based services prevalent in future Internet applications, including flexible methods for data classification and forwarding. This paper provides a brief description of representative platform hardware and software components. It then describes in more detail the feasibility of supporting a range of flexible and reconfigurable application-specific classification operations necessary for future Internet applications. Priyanka Tembey, Anish Bhatt, Subramanya Dulloor, Ada Gavrilovska, Karsten Schwan |
ICCCN | 4 |
| 2007 | Towards IQ-Appliances: Quality-awareness in Information VirtualizationabstractOur research addresses "information appliances' used in modern large-scale distributed systems to: (1) virtualize their data flows by applying actions such as filtering, format translation, etc., and (2) separate such actions from enterprise applications' business logic, to make it easier for future service-oriented codes to inter-operate in diverse and dynamic environments. Our specific contribution is the enrichment of runtimes of these appliances with methods for QoS-awareness, thereby giving them the ability to deliver desired levels of QoS even under sudden requirement changes - IQ-appliances. For experimental evaluation, we prototype an IQ-appliance. Measurements demonstrate the feasibility and utility of the approach. Radhika Niranjan, Ada Gavrilovska, Karsten Schwan, Priyanka Tembey |
NCA | 2 |
| 2007 | Advanced networking services for distributed multimedia streaming applications
Ada Gavrilovska, Srikanth Sundaragopalan, Karsten Schwan |
Multim. Tools Appl. | 1 |
| 2006 | Utilizing Network Processors in Distributed Enterprise EnvironmentsabstractThe integration of legacy systems and enterprise applications with novel communications, Internet, or networking services creates the problems of mismatches, information integration and possibly, problems from the evolution of the systems themselves. Enterprise services to solve these problems are currently implemented via commodity server hardware or application specific integrated circuits (ASICs), which suffer from either poor performance or a lack of flexibility. This paper presents an alternative to implementing these services by using network processors (NPs) as fast, flexible, and cost-efficient network appliances. It characterizes the hardware and software requirements of NP-based enterprise services, and presents a case study of a real-world enterprise problem. We implement a solution service to the problem on the Intel IXP2400, and present evaluation results that generalize the strengths, weaknesses, and requirements of the NP's role as a platform for enterprise service deployment Paul Royal, Mitch Halpin, Ada Gavrilovska, Karsten Schwan |
NCA | 3 |
| 2005 | Addressing data compatibility on programmable network platformsabstractLarge-scale applications require the efficient exchange of data across their distributed components, including data from heterogeneous sources and to widely varying clients. Inherent to such data exchanges are (1) discrepancies among the data representations used by sources, clients, or intermediate application components (e.g., due to natural mismatches or due to dynamic component evolution), and (2) requirements to route, combine, or otherwise manipulate data as it is being transferred. As a result, there is an ever growing need for data conversion services, handled by stubs in application servers, by middleware or messaging services, by the operating system, or by the network. This paper's goal is to demonstrate and evaluate the ability of modern network processors to efficiently address data compatibility issues, when data is 'in transit' between application-level services. Toward this end, we present the design and implementation of a network-level execution environment that permits systems to dynamically deploy and configure application-level data conversion services 'into' the network infrastructure. Experimental results obtained with a prototype implementation on Intel's IXP2400 network processors include measurements of XML-like data format conversions implemented with efficient binary data formats. Ada Gavrilovska, Karsten Schwan |
ANCS | 1 |
| 2005 | C-CORE: Using Communication Cores for High Performance Network ServicesabstractRecent hardware advances are creating multi-core systems with heterogeneous functionality. This paper explores how applications and middleware can utilize systems comprised of processors specialized for communication vs. computational tasks. The C-CORE execution environment enables applications, through middleware and underlying system functionality, to utilize both the computational capabilities of general purpose CPUs and the high performance communication hardware provided by specialized communication processors. Such future heterogeneous multi-core hardware is emulated by attaching a representative network processor - Intel's IXP2400 processor - to a general purpose CPU via a dedicated interconnect. For this platform, C-CORE provides abstractions to represent an application's communication actions, to efficiently couple such actions with application-level computations, and to dynamically create and configure the platform-resident 'chains' of computational and communication actions used by applications. C-CORE's functionality is evaluated with representative, communication-intensive applications. Measurements on our experimental platform establish the performance advantages afforded to applications by C-CORE Ada Gavrilovska, Karsten Schwan, Srikanth Sundaragopalan |
NCA | 2 |
| 2005 | Platform Overlays: enabling in-network stream processing in large-scale distributed applicationsabstractThe purpose of this research is to explore the capabilities of future, multi-core heterogeneous systems, with specialized communication support, to be used as efficient and flexible execution platforms in distributed streaming applications. On such platforms, we create overlays of hardware- and software-supported execution contexts -- platform overlays. Stream manipulations, represented via stream handlers, are deployed on top of such overlays, based on the ability of individual contexts to perform handler operations. As a result, stream processing is dynamically mapped to those platform resources best suited for it, and it can even be fully contained to the networking subsystems, thereby enabling in-network stream processing. Experimental results demonstrate the benefits of our approach towards meeting application-specific quality requirements. Ada Gavrilovska, Srikanth Sundaragopalan, Karsten Schwan |
NOSSDAV | 1 |
| 2002 | A Practical Approach for ?Zero? Downtime in an Operational Information SystemabstractAn operational information system (OIS) supports a real-time view of an organization's information critical to its logistical business operations. A central component of an OIS is an engine that integrates data events captured from distributed, remote sources in order to derive meaningful real-time views of current operations. This event derivation engine (EDE) continuously updates these views and also publishes them to a potentially large number of remote subscribers. The paper first describes a sample OIS and EDE in the context of an airline's operations. It then defines the performance and availability requirements to be met by this system, specifically focusing on the EDE component. One particular requirement for the EDE is that subscribers to its output events should not experience downtime due to EDE failures, crashes or increased processing loads. Toward this end, we develop and evaluate a practical technique for masking failures and for hiding the costs of recovery from EDE subscribers. This technique utilizes redundant EDEs that coordinate view replicas with a relaxed synchronous fault tolerance protocol. A combination of pre- and post-buffering of replicas is used to attain a solution that offers low response times (i.e., 'zero' downtime) while also preventing system failures in the presence of deterministic faults like 'ill-formed' messages. Parallelism realized via a cluster machine and application-specific techniques for reducing synchronization across replicas are used to scale a 'zero' downtime EDE to support the large number of subscribers it must service. Ada Gavrilovska, Karsten Schwan, Van Oleson |
ICDCS | 1 |
| 2001 | Adaptable Mirroring in Cluster ServersabstractThis paper presents a software architecture for continuously mirroring streaming data received by one node of a cluster-based server to other cluster nodes. The intent is to distribute the load on the server generated by the data's processing and distribution to many clients. This is particularly important when the server not only processes streaming data, but also performs additional processing tasks that heavily depend on current application state. One such task is the preparation of suitable initialization state for thin clients, so that such clients can understand future data events being streamed to them. In particular, when large numbers of thin clients must be initialized at the same time, initialization must be performed without jeopardizing the quality of service offered to regular clients continuing to receive data streams. The mirroring framework presented and evaluated has several novel aspects. First, by performing mirroring at the middleware level, application semantics may be used to reduce mirroring traffic, including filtering events based on their content, by coalescing certain events, or by simply varying mirroring rates according to current application needs concerning the consistencies of mirrored vs. original data. Second, we present an adaptive algorithm that varies mirror consistency and thereby, mirroring overheads in response to changes in clients' request behavior. Third, our framework not only mirrors events, but it can also mirror the new states computed from incoming events, thus enabling dynamic tradeoffs in the communication vs. computation loads imposed on the server node receiving events and on its mirror nodes. Ada Gavrilovska, Karsten Schwan, Van Oleson |
HPDC | 1 |