Christopher Stewart

dblp:99/1798 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-2860-7889ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 4 first-author · 10 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Poster: An Edge-to-Cloud Framework for Vigilance-Adaptive Drones
abstract
We present an edge-computing framework for autonomous drones to evaluate and adapt their missions in real time to minimize vigilant behaviors in monitored wildlife. Autonomous drones are increasingly valuable in ecology for capturing high-resolution videos of individual and group behaviors, yet their presence can bias data by provoking vigilance, i.e., heightened awareness triggered by perceived threats. To study this effect, we analyze the KABR (Kenyan Animal Behavior Recognition) dataset, where vigilance is indirectly measured through the frequency and duration of behaviors indicating induced vigilance. We found that vigilant behaviors varied based on environment, flight context, drone hardware, and species of interest, motivating an intelligent edge-to-cloud framework to evaluate the likelihood vigilance in real time.
Penelope Covey, Jenna Kline, Christopher Stewart
SEC3
2025 Edge-Native, Behavior-Adaptive Drone System for Wildlife Monitoring
abstract
Wildlife monitoring with drones must balance competing demands: approaching close enough to capture behaviorally-relevant video while avoiding stress responses that compromise animal welfare and data validity. Human operators face a fundamental attentional bottleneck: they cannot simultaneously control drone operations and monitor vigilance states across entire animal groups. By the time elevated vigilance becomes obvious, an adverse flee response by the animals may be unavoidable. To solve this challenge, we present an edge-native, behavior-adaptive drone system for wildlife monitoring. This configurable decision-support system augments operator expertise with automated group-level vigilance monitoring. Our system continuously tracks individual behaviors using YOLOv11m detection and YOLO-Behavior classification, aggregates vigilance states into a real-time group stress metric, and provides graduated alerts (alert vigilance → flee response) with operator-tunable thresholds for context-specific calibration. We derive service-level objectives (SLOs) from video frame rates and behavioral dynamics: to monitor 30fps video streams in real-time, our system must complete detection and classification within 33ms per frame. Our edge-native pipeline achieves 23.8ms total inference on GPU-accelerated hardware, meeting this constraint with a substantial margin. Retrospective analysis of seven wildlife monitoring missions demonstrates detection capability and quantifies the cost of reactive control: manual piloting results in 14 seconds average adverse behavior duration with 71.9% usable frames. Our analysis reveals operators could have received actionable alerts 51s before animals fled in 57% of missions. Simulating 5-second operator intervention yields a projected performance of 82.8% usable frames with 1-second adverse behavior duration, a 93% reduction compared to manual piloting.
Jenna Kline, Rugved Katole, Tanya Y. Berger-Wolf, Christopher Stewart
SEC4
2025 Poster: An Edge-Native Approach to Behavior-Adaptive Navigation in Drone Systems
abstract
Unmanned aerial vehicles, i.e., drones, are well-suited to monitor wildlife behaviors in remote habitats as they can quickly traverse rough terrain inaccessible to humans. However, drone operators must balance (1) potentially aggressive flight tactics to adequately monitor species of interest against (2) animal welfare and the stress that the drone may induce. The navigation strategy must adapt during flight in real-time based on the animals' reactions, while abiding by resource constraints imposed by the environment. We propose an edge-native framework that infers wildlife behaviors during flight and adapts navigation to consider animal welfare. When behavior-adaptive flight (BAF) miscues, the framework engages human-on-the-loop (HoTL) supervision, a light-weight solution to integrate expert feedback. We have validated our approach with high-fidelity digital twin simulations and real-world deployments. Our edge-native system integrates YOLO-v11m object detection and YOLO-Behavior action recognition, operating at 23.8ms inference latency while maintaining strict animal welfare thresholds.
Jenna Kline, Rugved Katole, Christopher Stewart
SEC3
2025 Poster: Managing Heterogeneity in Far-Edge AI
abstract
This poster paper examines the operating challenges that prevent far-edge AI systems from scaling, using Kubernetes-based agricultural drone deployments as a case study. Far-edge AI systems perform complex inference on data collected directly from the physical world, transforming fields from ecology to agronomy. However, these systems face diverse operating contexts including strict network firewalls, intermittent outages, and fully offline processing. We introduce inflection points, moments where edge site heterogeneity unexpectedly trigger failures, increase latency, and degrade AI inference. We propose research on inflection points that would allow container platforms to sense local operating context and adapt at runtime to ensure reliable execution at far-edge sites.
Vedant Patil, Christopher Stewart
SEC2
2025 Dataset assembly for training Spiking Neural Networks
Anthony Baietto, Christopher Stewart, Trevor J. Bihl
Neurocomputing2
2024 Domain-Aware Model Training as a Service for Use-Inspired Models
abstract
Use-inspired artificial intelligence (AI) tailors deep-learning models for image processing tasks in targeted scientific domains. These use-inspired models meet domain requirements for accuracy while parsimoniously using compact and efficient model architectures needed for inference in the field. However, before settling upon a model, domain experts repeatedly train and test models over a wide range of hyperparameters, contextual settings, and data configurations, making model training bespoke, time-consuming, and costly. Model-training-as-a-Service (MTaaS), i.e., cloud services designed for generic training workloads, can reduce training costs, but domain-aware designs and runtime adaptations could yield further reductions. This paper characterizes the potential for domain-aware design and runtime adaptation for MTaaS in digital agriculture. First, we studied the time to train models for 10 use-inspired agricultural datasets using pre-trained model weights derived from other agricultural datasets versus pre-trained weights derived from ImageNet, a widely used benchmark. Using agricultural datasets sped training time by up to 2X for some datasets, but provided modest speedups (<1.07) in the common case; Choosing the right dataset is critical. Next, we present an approach to predict training time given domain-aware pre-trained weights. Our predictions are strongly correlated with training time (r=0.93). Finally, we studied the use of domain-aware pre-trained weights in a MTaaS under Poisson and bursty arrival patterns for training tasks. Under bursty arrivals and tight memory constraints, domain-aware MTaaS reduced training time by 2.8 X and 12.2 X compared to model training using pre-trained ImageNet weights and from scratch, respectively.
Zichen Zhang 0004, Christopher Stewart
IC2E2
2024 Characterizing and Modeling AI-Driven Animal Ecology Studies at the Edge
abstract
Platforms that run artificial intelligence (AI) pipelines on edge computing resources are transforming the fields of animal ecology and biodiversity, enabling novel wildlife studies in animals' natural habitats. With emerging remote sensing hardware, e.g., camera traps and drones, and sophisticated AI models in situ, edge computing will be more significant in future AI-driven animal ecology (ADAE) studies. However, the study's objectives, the species of interest, its behaviors, range, and habitat, and camera placement affect the demand for edge resources at runtime. If edge resources are under-provisioned, studies can miss opportunities to adapt the settings of camera traps and drones to improve the quality and relevance of captured data. This paper presents salient features of ADAE studies that can be used to model latency, throughput objectives, and provision edge resources. Drawing from studies that span over fifty animal species, four geographic locations, and multiple remote sensing methods, we characterized common patterns in ADAE studies, revealing increasingly complex workflows involving various computer vision tasks with strict service level objectives (SLO). ADAE workflow demands will soon exceed individual edge devices' compute and memory resources, requiring multiple networked edge devices to meet performance demands. We developed a framework to scale traces from prior studies and replay them offline on representative edge platforms, allowing us to capture throughput and latency data across edge configurations. We used the data to calibrate queuing and machine learning models that predict performance on unseen edge configurations, achieving errors as low as 19%.
Jenna Kline, Austin O'Quinn, Tanya Y. Berger-Wolf, Christopher Stewart
SEC4
2024 Exploring nonintrusive measurements of spatio-temporal portrait of microservices
abstract
Abstract As cloud native technology advances, the scale and complexity of applications built on microservice architecture continue to expand, leading to increasingly intricate differences between software within the same application. Microservice applications, offering high flexibility, are deployed in data centers as black boxes from the users' perspective, leaving them with no insight into the orchestration of cloud service providers. Consequently, users face challenges in promptly recognizing performance imbalances within their deployed applications. Meanwhile, cloud service providers may cut costs by offering a mix of qualified and unqualified services, potentially deceiving users. To enhance the understanding of microservice application organization, we propose a non‐intrusive measurement framework, termed NMPI. NMPI facilitates rapid identification of microservice application defects, offering insights into cloud services and detecting fraudulent behavior in microservice‐based applications. We model microservice applications using a queue analysis‐based approach and filter the dominant frequency components of average response time signals by employing k‐means on the fast fourier transform (FFT). Our model constructs a library of performance portraits for various software, with these portraits resembling human fingerprints that carry and mark the software's internal information. Utilizing a two‐tier microservices‐based application incorporating a database as a case study allows us to demonstrate the effectiveness of NMPI. Our experimental results show that NMPI can produce differentiable profiles of data service performance portraits across a diverse and extensive range of workloads, enabling the identification of software types and the analysis of performance conditions.
Zichen Xu 0001, Dan Wu 0010, Xiaoling Li 0002, Biyong Liu, Haichuan Hu, Shuang Tan, Yusong Tan, Chenren Xu, Christopher Stewart, Qihe Zhou
Softw. Pract. Exp.10
2023 Poster: Profiling Edge Resource Demands of Zoom Maneuvers for Autonomous Unmanned Aerial Vehicles
Kevyn Angueira Irizarry, Christopher Stewart
SEC2
2023 Cost-Effective Strong Consistency on Scalable Geo-Diverse Data Replicas
abstract
The Raft algorithm maintains strong consistency across data replicas in Cloud. This algorithm places nodes, i.e., leader and follower, to serve read/write requests spanning geo-diverse sites. As the workload increases, Raft shall provide proportional scale-out performance. However, traditional scale-out techniques are bottlenecked in Raft with an exponentially increased performance penalty when provisioned sites exhaust local resources. To provide scalability in Raft, this paper presents a cost-effective mechanism that enables elastic auto scaling in Raft, called BW-Raft. BW-Raft extends the original Raft with the following abstractions: (1)secretarynodes that take over expensive log synchronization operations from the leader, relaxing the performance constraint on locks. (2)observernodes that handle reads only, improving throughput for typical data intensive services. These abstractions are stateless, allowing elastic scale-out on unreliable yet cheap spot instances. In theory, we prove that BW-Raft can preserve the strong consistency guarantee from Raft at scale-out, handling 50X more nodes, compared to the original Raft. We have prototyped the BW-Raft on key-value services and evaluated it with many state-of-the-arts on Amazon EC2 and Alibaba Cloud. Our results show that within the same budget, BW-Raft incurs 5-7X less resource footprint increment than Multi-Raft. Using spot instances, BW-Raft can reduces costs by 84.5%, compared to Multi-Raft. In the real world experiments, BW-Raft improves goodput of the 95th-percentile SLO by 9X, thus, serves an alternative for distributed service scaling out with strong consistency.
Yunxiao Du, Zichen Xu 0001, Kanqi Zhang, Christopher Stewart, Jiacheng Huang 0002
IEEE Trans. Cloud Comput.5
2022 Performance Modeling for Short-Term Cache Allocation
abstract
Short-term cache allocation grants and then revokes access to processor cache lines dynamically. For online services, short-term allocation can speed up targeted query executions and free up cache lines reserved, but normally not needed, for performance. However, in collocated settings, short-term allocation can increase cache contention, slowing down collocated query executions. To offset slowdowns, collocated services may request short-term allocation more often, making the problem worse. Short-term allocation policies manage which queries receive cache allocations and when. In collocated settings, these policies should balance targeted query speedups against slowdowns caused by recurring cache contention. We present a model-driven approach that (1) predicts response time under a given policy, (2) explores competing policies and (3) chooses policies that yield low response time for all collocated services. Our approach profiles cache usage offline, characterizes the effects of cache allocation policies using deep learning techniques and devises novel performance models for short-term allocation with online services. We tested our approach using data processing, cloud, and high-performance computing benchmarks collocated on Intel processors equipped with Cache Allocation Technology. Our models predicted median response time with 11% absolute percent error. Short-term allocation policies found using our approach out performed state-of-the-art shared cache allocation policies by 1.2–2.3X.
Christopher Stewart, Nathaniel Morris, Lydia Y. Chen, Robert Birke
ICPP1
2022 MARbLE: Multi-Agent Reinforcement Learning at the Edge for Digital Agriculture
abstract
Digital agriculture, hailed as the fourth great agricultural revolution, employs software-driven autonomous agents for in-field crop management. Edge computing resources deployed near crop fields support autonomous agents with substantial computational needs for tasks such as AI inference. In large fields, using multiple autonomous agents, called swarms, can speed up crop management tasks if sufficient edge resources are provisioned. However, to use swarms today, farmers and software developers craft their own standalone solutions that are either simple and ineffective or complicated and hard-to-reproduce. We present MARbLE, a platform for developing and managing swarms. MARbLE provides an easy-to-use programming paradigm that helps users build swarm workloads using multi-agent reinforcement learning. Developers supply just two functions Map() and Eval(). The platform automatically compiles and deploys swarms and continuously updates the reinforcement learning models that govern their actions. Developers can experiment with multiple swarm and edge resource configurations both in simulation and with actual in-field runs. We studied real UAV swarms conducting digital agriculture missions. We observe that swarms demanded edge computing resources in bursts; the ratio of average to peak demand was 2.9X. MARbLE uses energy-saving load balancing policies to duty cycle machines during workload demand troughs, leveraging workload patterns to save edge energy. Using MARbLE, we found that four-agent swarms with load balancing techniques sped up missions by 2.1X and reduced edge energy usage by up to 2X compared to state of the art autonomous swarms.
Jayson G. Boubin, Codi Burley, Peida Han, Barry Porter, Christopher Stewart
SEC6
2022 Bolt: Fast Inference for Random Forests
abstract
Random forests use ensembles of decision trees to boost accuracy for machine learning tasks. However, large ensembles slow down inference on platforms that process each tree in an ensemble individually. We present Bolt, a platform that restructures whole random forests, not just individual trees, to speed up inference. Conceptually, Bolt maps every path in each tree to a lookup table which, if cache were large enough, would allow inference with just one memory access. When the size of the lookup table exceeds cache capacity, Bolt employs a novel combination of lossless compression, parameter selection, and bloom filters to shrink the table while preserving fast inference. We compared inference speed in Bolt to three state-of-the-art platforms: Python Scikit-Learn, Ranger, and Forest Packing. We evaluated these platforms using datasets with vision, natural language processing and categorical applications. We observed that on ensembles of shallow decision trees Bolt can run 2--14X faster than competing platforms and that Bolt's speedups persist as the number of decision trees in an ensemble increases.
Eduardo Romero 0005, Christopher Stewart, Kyle C. Hale, Nathaniel Morris
Middleware2
2022 Adaptive Deployment for Autonomous Agricultural UAV Swarms
abstract
Unmanned aerial vehicles (UAV) play a critical role in many edge computing deployments and applications. UAV are prized for their maneuverability, low cost, and sensing capacity, facilitating many applications that would otherwise be prohibitively expensive or dangerous without them. UAV are cheaper than alternative aerial analysis methods, but still incur costs from expensive human piloting and workloads which necessitate high-resolution coverage of large areas. Recently, autonomous UAV swarms have emerged to increase the speed of deployments, decrease the cost and scope of human piloting, and improve the quality of autonomous decision-making through data sharing. Autonomous UAV deployments, however, suffer from external factors. UAV are inherently power-constrained, with low onboard battery lives and limited ability to siphon power from the edge systems that support them. Certain environmental conditions, like inclement weather, wind, extreme heat, and low light also affect UAV power consumption, sensed data quality, and ultimately mission success. In this paper, we present an empirically based model for efficient autonomous swarm deployment. We built and deployed a real autonomous UAV swarm to map leaf defoliation in soybeans. Using this deployment, we determined environmental conditions which led to malfunctions, inefficient edge energy usage, and mispredictions. Using these findings, we developed a deployment model for UAV swarms that decreases malfunctions and data irregularities by 4.9X and decreases edge energy consumption by 45%, while increasing deployment times by only 4%.
Jayson G. Boubin, Zichen Zhang 0004, John Chumley, Christopher Stewart
SenSys4
2021 Avis: In-Situ Model Checking for Unmanned Aerial Vehicles
abstract
Control firmware in unmanned aerial vehicles (UAVs) uses sensors to model and manage flight operations, from takeoff to landing to flying between waypoints. However, sensors can fail at any time during a flight. If control firmware mishandles sensor failures, UAVs can crash, fly away, or suffer other unsafe conditions. In-situ model checking finds sensor failures that could lead to unsafe conditions by systematically failing sensors. However, the type of sensor failure and its timing within a flight affect its manifestation, creating a large search space. We propose Avis, an in-situ model checker to quickly uncover UAV sensor failures that lead to unsafe conditions. Avis exploits operating modes, i.e., a label that maps software execution to corresponding flight operations. Widely used control firmware already support operating modes. Avis injects sensor failures as the control firmware transitions between modes - a key execution point where mishandled software exceptions can trigger unsafe conditions. We implemented Avis and applied it to ArduPilot and PX4. Avis found unsafe conditions 2.4X faster than Bayesian Fault Injection, the leading, state-of-theart approach. Within the current code base of ArduPilot and PX4, Avis discovered 10 previously unknown software bugs that lead to unsafe conditions. Additionally, we reinserted 5 known bugs that caused serious, unsafe conditions and Avis correctly reported all of them.
Max Taylor, Haicheng Chen, Christopher Stewart
DSN4
2020 Poster: Configuration Management for Internet Services at the Edge: A Data-Driven Approach
abstract
Internet services are increasingly pushed from the remote cloud to the edge sites close to data sources to offer fast response time and low energy footprint. However, software deployed at edge sites must be updated frequently. Performing updates as soon as they are available consumes a large amount of energy. Configuration management tools that install software updates and manage allowed staleness can inflate energy demands, especially when updates interrupt idle periods at the edge site and block processors from entering power-saving modes. Our research studies configuration management policies, their effect on energy footprint and strategies to optimize them. We have observed that policies yielding low energy footprint differ from site to site and over time. We propose a data-driven approach that uses data collected at each edge site to predict an energy-efficient policy and also guards against worst-case performance if data-driven predictions error occurs. We use a novel randomwalk approach to manage data-driven policies that yield a low footprint for a representative trace of updates observed at an edge site. We are setting up 4 edge service benchmarks powered by AI inference to create realistic software update traces.
Christopher Stewart
SEC2
2019 Elastic, geo-distributed RAFT
abstract
Raft is a protocol to maintain strong consistency across data replicas in cloud. It is widely used, especially by workloads that span geographically distributed sites. As these workloads grow, Raft's costs should grow, as least proportionally. However, auto scaling approaches for Raft inflate costs by provisioning at all sites when one site exhausts its local resources. This paper presents Geo-Raft, a scale-out mechanism that enables precise auto scaling for Raft. Geo-Raft extends Raft with the following abstractions: (1) secretaries which takes log processing for the leader and (2) observers which process read requests for followers. These abstractions are stateless, allowing for elastic auto scaling, even on unreliable spot instances. Geo-Raft provably preserves strong consistency guarantees provided by Raft. We implemented and evaluated Geo-Raft with multiple auto scaling techniques on Amazon EC2. Geo-Raft scales in resource footprint increments 5-7X smaller than Multi-Raft, the state of the art. Using spot instances, Geo-Raft reduces costs by 84.5% compared to Multi-Raft. Geo-Raft improves goodput of 95th-percentile SLO by 9X.Geo-Raft operates key-value services for 6 months without losing data or crash.
Zichen Xu 0001, Christopher Stewart, Jiacheng Huang 0002
IWQoS2
2018 SLO Computational Sprinting
abstract
No abstract available.
Nathaniel Morris, Indrajeet Saravanan, Pollyanna Cao, Jerry Ding, Christopher Stewart
SoCC5
2018 Model-driven computational sprinting
abstract
Computational sprinting speeds up query execution by increasing power usage for short bursts. Sprinting policy decides when and how long to sprint. Poor policies inflate response time significantly. We propose a model-driven approach that chooses between sprinting policies based on their expected response time. However, sprinting alters query executions at runtime, creating a complex dependency between queuing and processing time. Our performance modeling approach employs offline profiling, machine learning, and first-principles simulation. Collectively, these modeling techniques capture the effects of sprinting on response time. We validated our modeling approach with 3 sprinting mechanisms across 9 workloads. Our performance modeling approach predicted response time with median error below 4% in most tests and median error of 11% in the worst case. We demonstrated model-driven sprinting for cloud providers seeking to colocate multiple workloads on AWS Burstable Instances while meeting service level objectives. Model-driven sprinting uncovered policies that achieved response time goals, allowing more workloads to colocate on a node. Compared to AWS Burstable policies, our approach increased revenue per node by 1.6X.
Nathaniel Morris, Christopher Stewart, Lydia Y. Chen, Robert Birke, Jaimie Kelley
EuroSys2
2018 Assessing Learning Behavior and Cognitive Bias from Web Logs
abstract
This research to practice, work in progress paper presents the analysis strategy used to assess the learning behavior using logs on an e-learning platform. Students who can link algebraic functions to their corresponding graphs perform well in STEM courses. Early algebra curricula teaches these concepts in tandem. However, it is challenging to assess whether students are linking the concepts. Video analyses, interviews and other traditional methods that aim to quantify how students link the concepts taught in school require precious classroom and teacher time. We use web logs to infer learning. Web logs are widely available and amenable to data science. Our approach partitions the web interface into components related to data and graph concepts. We collect click and mouse movement data as users interact with these components. We used statistical and data mining techniques to model their learning behavior. We built our models to assess learning behavior for a workshop presented in Summer 2016. Students in the workshop were middle-school math teachers planning to use this curriculum in their own classrooms. We used our models to assess participation levels, a prerequisite indicator for learning. Our models aligned with ground-truth traditional methods for 17 of 18 students. The results of the models with respect to the two types of components of the web portal have been used to infer possible data or graph oriented cognitive bias.
Rashmi Rao, Christopher Stewart, Arnulfo Pérez, Siva Meenakshi Renganathan
FIE2
2017 Early work on modeling computational sprinting
abstract
Ever tightening power caps constrain the sustained processing speed of modern processors. With computational sprinting, processors reserve a small power budget that can be used to increase processing speed for short bursts. Computational sprinting speeds up query executions that would otherwise yield slow response time. Common mechanisms used for sprinting include DVFS, core scaling, CPU throttling and application-specific accelerators.
Nathaniel Morris, Christopher Stewart, Robert Birke, Lydia Y. Chen, Jaimie Kelley
SoCC2
2017 Preliminary results on an interactive learning tool for early algebra education
abstract
Interactive learning tools allow students to explore STEM concepts deeply, improving educational outcomes. Interactive tools that use cloud computing resources can (1) explore computationally intensive concepts and (2) seamlessly link concepts to graphical representations. When these goals conflict, good design, and implementation principles are needed to (1) preserve interactive response times and (2) uphold curriculum goals. For this paper, we present an interactive cloud-based learning tool that allows students to interactively explore algebraic formulas, their corresponding graphs (including axes) and generating data. Our tool is designed with both curriculum and interactive systems support in mind and it integrates with existing classroom management platforms. Our implementation employs a careful division of work between client-side browsers and cloud servers to provide instant response times that are within human perception limits. We achieve these design goals by (1) client-side programming for interactive components, (2) sending student activity data to servers at different rates for different sharing types, and (3) align student's data with Moodle's (a popular open source classroom management system) schema. Twenty Math teachers were trained to teach Algebra to students using our interactive service. Over eighty percent of teachers trained with our system adopted it. Now, our system is deployed in 5 schools, 30 classrooms, and 1400 students.
Siva Meenakshi Renganathan, Christopher Stewart, Arnulfo Pérez, Rashmi Rao, Bailey Braaten
FIE2
2017 A Pareto Framework for Data Analytics on Heterogeneous Systems: Implications for Green Energy Usage and Performance
abstract
Distributed algorithms for data analytics partition their input data across many machines for parallel execution. At scale, it is likely that some machines will perform worse than others because they are slower, power constrained or dependent on undesirable, dirty energy sources. It is challenging to balance analytics workloads across heterogeneous machines because the algorithms are sensitive to statistical skew in data partitions. A skewed partition can slow down the whole workload or degrade the quality of results. Sizing partitions in proportion to each machine's performance may introduce or further exacerbate skew. In this paper, we propose a scheme that controls the statistical distribution of each partition and sizes partitions according to the heterogeneity of the computing environment. We model heterogeneity as a multi-objective optimization, with the objectives being functions for execution time and dirty energy consumption. We use stratification to control skew. Experiments show that our computational heterogeneity-aware (Het-Aware) partitioning strategy speeds up running time by up to 51% over the stratified partitioning scheme baseline. We also have a heterogeneity and energy aware (Het-Energy-Aware) partitioning scheme which is slower than the Het-Aware solution but can lower the dirty energy footprint by up to 26%. For some analytic tasks, there is also a significant qualitative benefit when using such partitioning strategies.
Aniket Chakrabarti, Srinivasan Parthasarathy 0001, Christopher Stewart
ICPP3
2016 Blending on-demand and spot instances to lower costs for in-memory storage
abstract
In cloud computing, workloads that lease instances on demand get to execute exclusively for a set time. In contrast, workloads that lease spot instances execute until a competing workload outbids the current lease. Spot instances cost less than on-demand instances, but few workloads can use spot instances because of the variable leasing period. We present BOSS, a framework that uses spot instances to reduce costs for in-memory storage workloads. BOSS uses on-demand instances to create and update objects. It uses spot instances to handle read-only queries. BOSS leases instances from multiple sites and exploits varying prices between the sites. When spot instances stop abruptly at one site, BOSS places newly created objects at other sites, reducing the impact on response time. BOSS proposes a novel, online replication approach (1) avoids placing data at too many sites and (2) provides O(1.5)-competitive ratio under skewed cost distributions. Within a site, BOSS manages the tradeoff between savings and risks from replicating to spot instances. We implemented BOSS on top of Cassandra and deployed it on up to 78 instances across 8 sites in Amazon and Google clouds. With BOSS hosting TPC-W data, we spent $8 per hour on Amazon. For the same service, we spent $55 per hour to use ElastiCache and $49 per hour to use on-demand instances only. BOSS saved 85% and 84% respectively. Further, BOSS achieved 95th percentile response time within 13% of ElastiCache.
Zichen Xu 0001, Christopher Stewart
INFOCOM2
2014 Managing Tiny Tasks for Data-Parallel, Subsampling Workloads
abstract
Subsampling workloads compute statistics from a set of observed samples using a random subset of sample data (i.e., a subsample). Data-parallel platforms group these samples into tasks, each task subsamples its data in parallel. In this paper, we study subsampling workloads that benefit from tiny tasks-i.e., tasks comprising few samples. Tiny tasks reduce processor cache misses caused by random subsampling, which speeds up per-task running time. However, they can also cause significant scheduling overheads that negate the time reduction from reduced cache misses. For example, vanilla Hadoop takes longer to start tiny tasks than to run them. We compared the task scheduling overheads of vanilla Hadoop, lightweight Hadoop setups, and BashReduce. BashReduce, the best platform, outperformed the worst by 3.6X but scheduling overhead was still 12% of a task's running time. We improved BashReduce's scheduler by allowing it to size tasks according to kneepoints on the miss rate curve. We tested these changes on high-throughput genotype data and on data obtained from Netflix. Our improved BashReduce outperformed vanilla Hadoop by almost 3X and completed short, interactive jobs almost as efficiently as long jobs. These results held at scale and across diverse, heterogeneous hardware.
Sundeep Kambhampati, Jaimie Kelley, Christopher Stewart, William C. L. Stewart, Rajiv Ramnath
IC2E3
2012 Policy and mechanism for carbon-aware cloud applications
abstract
This position paper explores research challenges facing carbon-aware cloud applications. These applications run inside of a renewable-energy datacenter, provision resources on demand, and seek to minimize their use of carbon-heavy, grid energy. First, we argue that carbon-aware applications need new provisioning policies that address the uncertainty of renewable energy. Second, we argue that renewable-energy datacenters need mechanisms to determine the contribution of grid energy to specific application workloads. We propose first-cut solutions to these problems and present preliminary results for carbon-aware Web server running on a small renewable-energy cluster.
Christopher Stewart, Daniel Gmach, Martin F. Arlitt
NOMS2
2010 EntomoModel: Understanding and Avoiding Performance Anomaly Manifestations
abstract
Subtle implementation errors or mis-configurations in complex Internet services may lead to performance degradations without causing failures. These undiscovered performance anomalies afflict many of today's systems, causing violations of service-level agreements (SLAs), unnecessary resource over provisioning, or both. In this paper, we re-inserted realistic anomaly causes into a multi-tier Internet service architecture and studied their manifestations. We observed that each cause had certain workload and management parameters that were more likely to trigger manifestations, hinting that such parameters could be effective classifiers. This observation held even when anomaly causes manifested differently in combination than in isolation. Our study motivates EntomoModel, a framework for depicting performance anomaly manifestations. EntomoModel uses decision tree classification and a design-driven performance model to characterize the workload and management policy settings under which manifestations are likely. EntomoModel enables online system management that avoids anomaly manifestations by dynamically adjusting system management parameters. Our trace-driven evaluations show that manifestation avoidance based on EntomoModel, or entomophobic management, can reduce 98th percentile SLA violations by 67% compared to an anomaly oblivious adaptive approach. In a cloud computing scenario with elastic resource allocation, our approach uses less than half of the resources needed in static over-provisioning.
Christopher Stewart, Arun Iyengar, Jian Yin 0002
MASCOTS1
2008 Hardware counter driven on-the-fly request signatures
abstract
Today's processors provide a rich source of statistical informationon application execution through hardware counters. In this paper, we explore the utilization of these statistics as request signaturesin server applications for identifying requests and inferring high-level request properties (e.g., CPU and I/O resource needs). Our key finding is that effective request signatures may be constructed using a small amount of hardware statistics while the request is still in an early stage of its execution. Such on-the-fly request identification and property inference allow guided operating system adaptation at request granularity (e.g., resource-aware request scheduling and on-the-fly request classification). We address the challenges of selecting hardware counter metrics for signature construction and providing necessary operating system support for per-request statistics management. Our implementation in the Linux 2.6.10 kernel suggests that our approach requires low overhead suitable for runtime deployment. Our on-the-fly request resource consumption inference (averaging 7%, 3%, 20%, and 41% prediction errors for four server workloads, TPC-C, TPC-H, J2EE-based RUBiS, and a trace-driven index search, respectively) is much more accurate than the online running-average based prediction (73-82% errors). Its use for resource-aware request scheduling results in a 15-70% response time reduction for three CPU-bound applications. Its use for on-the-fly request classification and anomaly detection exhibits high accuracy for the TPC-H workload with synthetically generated anomalous requests following a typical SQL-injection attack pattern.
Ming Zhong 0006, Sandhya Dwarkadas, Chuanpeng Li, Christopher Stewart
ASPLOS5
2008 Operational Analysis of Parallel Servers
Terence Kelly, Alex Zhang, Christopher Stewart
MASCOTS4
2008 Operational analysis of processor speed scaling
abstract
This brief announcement presents a pair of performance laws that bound the change in aggregate job queueing time that results when the processor speed changes in a parallel computing system. Our laws require only lightweight passive external observations of a black-box system and they apply to many commonly employed scheduling policies. By predicting the application-level performance impact of processing speed adjustments in parallel processors, including traditional SMPs and now increasingly ubiquitous multicore processors, our laws address problems ranging from capacity planning to dynamic resource allocation. Finally, our results show that operational analysis---an approach to performance analysis traditionally associated with commercial transaction processing systems---usefully complements existing parallel performance analysis techniques.
Alex Zhang, Terence Kelly, Christopher Stewart
SPAA4
2008 A Dollar from 15 Cents: Cross-Platform Management for Internet Services
Christopher Stewart, Terence Kelly, Alex Zhang
USENIX ATC1
2007 Exploiting nonstationarity for performance prediction
abstract
Real production applications ranging from enterprise applications to large e-commerce sites share a crucial but seldom-noted characteristic: The relative frequencies of transaction types in their workloads are nonstationary, i.e., the transaction mix changes over time. Accurately predicting application-level performance in business-critical production applications is an increasingly important problem. However, transaction mix nonstationarity casts doubt on the practical usefulness of prediction methods that ignore this phenomenon.
Christopher Stewart, Terence Kelly, Alex Zhang
EuroSys1
2006 Realism and simplicity: disk simulation for instructional OS performance evaluation
abstract
Operating system laboratory assignments based on bare hardware or detailed machine simulators can be excessively challenging for many students. In the most often used approach, students develop kernels on virtual machines with a much simplified hardware interface. Traditionally this simplification goes so far as to make realistic performance measurement impossible. We propose Vesper, an instructional disk drive simulator with a high degree of performance realism. Vesper retains simplicity while providing timing statistics close to that of real disk drives. The key to our approach is to provide hardware abstractions that are simple but yet capable of capturing device interactions with major performance impacts. Vesper laboratory assignments allow students to realistically explore the performance consequences of various system designs without the cumbersome aspects of the real hardware interface. This paper describes the design and implementation of the Vesper disk drive simulator. We evaluate the effectiveness of Vesper-based laboratory assignments in terms of operating system performance evaluation. Student experience and feedback are also reported.
Peter DeRosa, Christopher Stewart, Jonathan Pearson
SIGCSE3
2005 Performance Modeling and System Management for Multi-component Online Services
Christopher Stewart
NSDI1
2005 Daphne: performance debugging using model-driven anomaly characterization
abstract
It is not uncommon for complex systems (like multi-component online services) to perform worse than expected. Configuration errors, buggy code, and tricky middleware/OS behavior induce performance anomalies under certain workloads and system configurations. It is challenging to pinpoint root causes of such performance problems in complex systems. Currently, performance problems are solved by system administrators who intimately understand the system and are willing to forgo a personal life.
Christopher Stewart
SOSP1