VLDB 2026 Research / reviewers in the wild / expert
Timothy Zhu
dblp:121/1838
· DBLP profile ↗
22ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0001-8394-8953ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 7 · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Workload Intelligence: Workload-Aware IaaS abstraction for Cloud EfficiencyabstractToday, cloud workloads are largely opaque to the cloud platform. Typically, the only information the platform receives is the virtual machine (VM) type and possibly a decoration to the type (e.g., the VM is evictable). Similarly, workloads receive minimal information from the platform; generally, only telemetry from their VMs or occasional signals (e.g., just before a VM is evicted). The narrow interface between workloads and platforms has several drawbacks: (1) a surge in VM types and decorations in public cloud platforms complicates customer selection; (2) key workload characteristics (e.g., low availability requirements) are often unspecified, hindering platform customization for optimized resource usage and cost savings; and (3) workloads may be unaware of potential optimizations or lack sufficient time to react to platform events. To resolve these issues and improve cloud efficiency, we propose Workload Intelligence (WI), a framework for enabling dynamic bi-directional communication between cloud workloads and cloud platform. Lexiang Huang, Anjaly Parayil, Xiaoting Qin, Chetan Bansal, Jovan Stojkovic, Pantea Zardoshti, Pulkit A. Misra, Eli Cortez, Raphael Ghelman, Íñigo Goiri, Saravan Rajmohan, Jim Kleewein, Rodrigo Fonseca, Timothy Zhu, Ricardo Bianchini |
SC | 15 |
| 2025 | Towards Workload-aware Cloud Efficiency: A Large-scale Empirical Study of Cloud Workload CharacteristicsabstractCloud providers introduce features and optimizations to improve efficiency and reliability, such as Spot VMs, Harvest VMs, oversubscription, and auto-scaling. To use these effectively, it's important to understand workload characteristics. However, workload characterization can be complex and difficult to scale manually due to multiple signals involved. In this study, we conduct the first large-scale empirical study of first-party workloads at Microsoft to understand their characteristics. Through this empirical study, we aim to answer the following questions: (1) What are the critical workload characteristics that impact efficiency and reliability on cloud platforms? (2) How do these characteristics vary across different workloads? (3) How can cloud platforms leverage these insights to efficiently characterize all workloads at scale? This study provides a deeper understanding of workload characteristics and their impact on cloud performance, which can aid in optimizing cloud services and identifies potential areas for future research. Anjaly Parayil, Xiaoting Qin, Íñigo Goiri, Lexiang Huang, Timothy Zhu, Chetan Bansal |
ICPE | 6 |
| 2025 | TraceScaler: A Framework for Scaling Load in Real-World Traces for System EvaluationabstractTrace replay is a common approach for evaluating systems by rerunning historical traffic patterns, but it’s not always possible to find suitable real-world traces at the desired level of system load. To experiment with different loads, one needs to downscale a trace to decrease the load or upscale a trace to artificially increase the load. This article expands upon our work, TraceUpscaler [ 92 ], by considering the interaction of upscaling and downscaling. In addition to evaluating upscaling with traces collected from a subset of the cluster, we also evaluate upscaling with traces that were downscaled with the state-of-the-art downscaling tool, TraceSplitter [ 91 ], to demonstrate that the upscaling and downscaling techniques are compatible and do not introduce unexpected artifacts in the scaling. In addition to comparing against prior approaches, we develop a novel upscaling technique, TraceOverlap , based on the idea of overlapping different time periods in a trace, where we identify the most similar time periods to overlap. Our evaluation demonstrates that TraceUpscaler and TraceOverlap are both more accurate in maintaining latency characteristics than prior approaches, with TraceUpscaler matching the original trace latency more closely. Finally, we provide a unified framework, TraceScaler , that combines TraceUpscaler with TraceSplitter to provide experimenters a common tool for their trace scaling needs. Sultan Mahmud Sajal, Salman Estyak, Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
ACM Trans. Comput. Syst. | 4 |
| 2024 | AutoBurst: Autoscaling Burstable Instances for Cost-effective Latency SLOsabstractBurstable instances provide a low-cost option for consumers using the public cloud, but they come with significant resource limitations. They can be viewed as "fractional instances" where one receives a fraction of the compute and memory capacity at a fraction of the cost of regular instances. The fractional compute is achieved via rate limiting, where a unique characteristic of the rate limiting is that it allows for the CPU to burst to 100% utilization for limited periods of time. Prior research has shown how this ability to burst can be used to serve specific roles such as a cache backup and handling flash crowds. Our work provides a general-purpose approach to meeting latency SLOs via this burst capability while optimizing for cost. AutoBurst is able to achieve this by controlling both the number of burstable and regular instances along with how/when they are used. Evaluations show that our system is able to reduce cost by up to 25% over the state-of-the-art while maintaining latency SLOs. Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar |
SoCC | 2 |
| 2024 | TraceUpscaler: Upscaling Traces to Evaluate Systems at High LoadabstractTrace replay is a common approach for evaluating systems by rerunning historical traffic patterns, but it is not always possible to find suitable real-world traces at the desired level of system load. Experimenting with higher traffic loads requires upscaling a trace to artificially increase the load. Unfortunately, most prior research has adopted ad-hoc approaches for upscaling, and there has not been a systematic study of how the upscaling approach impacts the results. One common approach is to count the arrivals in a predefined time-interval and multiply these counts by a factor, but this requires generating new requests/jobs according to some model (e.g., a Poisson process), which may not be realistic. Another common approach is to divide all the timestamps in the trace by an upscaling factor to squeeze the requests into a shorter time period. However, this can distort temporal patterns within the input trace. This paper evaluates the pros and cons of existing trace upscaling techniques and introduces a new approach, TraceUpscaler, that avoids the drawbacks of existing methods. The key idea behind TraceUpscaler is to decouple the arrival timestamps from the request parameters/data and upscale just the arrival timestamps in a way that preserves temporal patterns within the input trace. Our work applies to open-loop traffic where requests have arrival timestamps that aren't dependent on previous request completions. We evaluate TraceUpscaler under multiple experimental settings using both real-world and synthetic traces. Through our study, we identify the trace characteristics that affect the quality of upscaling in existing approaches and show how TraceUpscaler avoids these pitfalls. We also present a case study demonstrating how inaccurate trace upscaling can lead to incorrect conclusions about a system's ability to handle high load. Sultan Mahmud Sajal, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
EuroSys | 2 |
| 2024 | Fast and Accurate DNN Performance Estimation across Diverse Hardware PlatformsabstractPerformance modeling is an important tool for many purposes such as designing hardware accelerators, improving scheduling, optimizing system parameters, procuring new hardware, etc. This paper provides a new methodology for constructing performance models for Deep Neural Networks (DNNs), a popular machine learning workload. Prior works require running DNNs on existing hardware, which may not be available, or simulating the computation on futuristic hardware, which is slow and not scalable. We instead take an analytical approach based on analyzing the raw operations within DNN algorithms, which allows us to estimate performance across any hardware, even hardware that is in the process of being designed. Evaluations show our approach is fast and gives a good first order approximation ($\pm 10-15 \%$ accuracy) across many DNNs and hardware platforms including GPUs, CPUs, and a futuristic Processing In Memory (PIM) accelerator called BLIMP. Vishwas Vasudeva Kakrannaya, Siddhartha Balakrishna Rai, Anand Sivasubramaniam, Timothy Zhu |
MASCOTS | 4 |
| 2023 | Kerveros: Efficient and Scalable Cloud Admission Control
Sultan Mahmud Sajal, Luke Marshall, Beibin Li, Shandan Zhou, Abhisek Pan, Konstantina Mellou, Deepak Narayanan, Timothy Zhu, David Dion, Thomas Moscibroda, Ishai Menache |
OSDI | 8 |
| 2022 | Metastable Failures in the Wild
Lexiang Huang, Matt Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko |
OSDI | 7 |
| 2022 | Overflowing emerging neural network inference tasks from the GPU to the CPU on heterogeneous serversabstractWhile current deep learning (DL) inference runtime systems sequentially offload the model's tasks on to an available GPU/accelerator based on its capability, we make a case for selectively redirecting some of these tasks to the CPU and running them concurrently with the GPU doing other work. This new opportunity specifically arises for emerging DL models whose data flow graphs (DFGs) have much wider fan-outs compared to traditional ones which are invariably linear chains of tasks. By opportunistically moving some of these tasks to the CPU, we can (i) shave off service times from the critical path of the DFG, (ii) devote the GPU for more deserving tasks, and (iii) improve overall utilization of the provisioned hardware in the server. However, several factors such as its criticality in the DFG, slowdown when moved to a different hardware engine, and overheads in transferring input/output data across these engines, determine the what/when/how of tasks to be directed. While this is computationally demanding and slow to be solved optimally, through a series of rationales we derive a fast technique for task overflow from GPU to CPU. We implement this technique on a nimble heterogeneous concurrent runtime engine built on top of the state-of-the-art ONNXRuntime engine and demonstrate > 10% reduction in latency, > 19% gain in throughput, and > 9.8% savings in GPU memory usage for emerging neural network models. Adithya Kumar, Anand Sivasubramaniam, Timothy Zhu |
SYSTOR | 3 |
| 2021 | tprof: Performance profiling via structural aggregation and automated analysis of distributed systems tracesabstractThe traditional approach for performance debugging relies upon performance profilers (e.g., gprof, VTune) that provide average function runtime information. These aggregate statistics help identify slow regions affecting the entire workload, but they are ill-suited for identifying slow regions that only impact a fraction of the workload, such as tail latency effects. This paper takes a new approach to performance profiling by utilizing distributed tracing systems (e.g., Dapper, Zipkin, Jaeger). Since traces provide detailed timing information on a per-request basis, it is possible to group and aggregate tracing data in many different ways to identify the slow parts of the system. Our new approach to trace aggregation uses the structure embedded within traces to hierarchically group similar traces and calculate increasingly detailed aggregate statistics based on how the traces are grouped. We also develop an automated tool for analyzing the hierarchy of statistics to identify the most likely performance issues. Our case study across two complex distributed systems illustrates how our tool is able to find multiple performance issues that lead to 10x and 28x performance improvements in terms of average and tail latency, respectively. Our comparison with a state-of-the-art industry tool shows that our tool can pinpoint performance slowdowns more accurately than current approaches. Lexiang Huang, Timothy Zhu |
SoCC | 2 |
| 2021 | TraceSplitter: a new paradigm for downscaling tracesabstractRealistic experimentation is a key component of systems research and industry prototyping, but experimental clusters are often too small to replay the high traffic rates found in production traces. Thus, it is often necessary to downscale traces to lower their arrival rate, and researchers/practitioners generally do this in an ad-hoc manner. For example, one practice is to multiply all arrival timestamps in a trace by a scaling factor to spread the load across a longer timespan. However, temporal patterns are skewed by this approach, which may lead to inappropriate conclusions about some system properties (e.g., the agility of auto-scaling). Another popular approach is to count the number of arrivals in fixed-sized time intervals and scale it according to some modeling assumptions. However, such approaches can eliminate or exaggerate the fine-grained burstiness in the trace depending on the time interval length. Sultan Mahmud Sajal, Rubaba Hasan, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha Sen 0001 |
EuroSys | 3 |
| 2021 | Metastable failures in distributed systemsabstractWe describe metastable failures---a failure pattern in distributed systems. Currently, metastable failures manifest themselves as black swan events; they are outliers because nothing in the past points to their possibility, have a severe impact, and are much easier to explain in hindsight than to predict. Although instances of metastable failures can look different at the surface, deeper analysis shows that they can be understood within the same framework. Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, Timothy Zhu |
HotOS | 4 |
| 2020 | Peafowl: in-application CPU scheduling to reduce power consumption of in-memory key-value storesabstractThe traffic load sent to key-value (KV) stores varies over long timescales of hours to short timescales of a few microseconds. Long-term variations present the opportunity to save power during low or medium periods of utilization. Several techniques exist to save power in servers, including feedback-based controllers that right-size the number of allocated CPU cores, dynamic voltage and frequency scaling (DVFS), and c-state (idle-state) mechanisms. In this paper, we demonstrate that existing power saving techniques are not effective for KV stores. This is because the high rate of traffic even under low load prevents the system from entering low power states for extended periods of time. To achieve power savings, we must unbalance the load among the CPU cores so that some of them can enter low power states during periods of low load. We accomplish this by introducing the notion of in-application CPU scheduling. Instead of relying on the kernel to schedule threads, we pin threads to bypass the kernel CPU scheduler and then perform the scheduling within the KV store application. Our design, Peafowl, is a KV store that features an in-application CPU scheduler that monitors the load to learn the workload characteristics and then scales the number of active CPU cores when the load drops, leading to notable power savings during low or medium periods of utilization. Our experiments demonstrate that Peafowl uses up to 40--54% lower power than state of the art approaches such as Rubik and μDPM. Esmail Asyabi, Azer Bestavros, Erfan Sharafzadeh, Timothy Zhu |
SoCC | 4 |
| 2020 | The Fast and The Frugal: Tail Latency Aware Provisioning for Coping with Load VariationsabstractSmall and medium sized enterprises use the cloud for running online, user-facing, tail latency sensitive applications with well-defined fixed monthly budgets. For these applications, adequate system capacity must be provisioned to extract maximal performance despite the challenges of uncertainties in load and request-sizes. In this paper, we address the problem of capacity provisioning under fixed budget constraints with the goal of minimizing tail latency. Adithya Kumar, Iyswarya Narayanan, Timothy Zhu, Anand Sivasubramaniam |
WWW | 3 |
| 2019 | BurScale: Using Burstable Instances for Cost-Effective Autoscaling in the Public CloudabstractCloud providers have recently introduced burstable instances - virtual machines whose CPU capacity is rate limited by token-bucket mechanisms. A user of a burstable instance is able to burst to a much higher resource capacity ("peak rate") than the instance's long-term average capacity ("sustained rate"), provided the bursts are short and infrequent. A burstable instance tends to be much cheaper than a conventional instance that is always provisioned for the peak rate. Consequently, cloud providers advertise burstable instances as cost-effective options for customers with intermittent needs and small (e.g., single VM) clusters. By contrast, this paper presents two novel usage scenarios for burstable instances in larger clusters with sustained usage. We demonstrate (i) how burstable instances can be utilized alongside conventional instances to handle the transient queueing arising from variability in traffic, and (ii) how burstable instances can mask the VM startup/warmup time when autoscaling to handle flash crowds. We implement our ideas in a system called BurScale and use it to demonstrate cost-effective autoscaling for two important workloads: (i) a stateless web server cluster, and (ii) a stateful Memcached caching cluster. Results from our prototype system show that via its careful combination of burstable and regular instances, BurScale can ensure similar application performance as traditional autoscaling systems that use all regular instances while reducing cost by up to 50%. Ataollah Fatahi Baarzi, Timothy Zhu, Bhuvan Urgaonkar |
SoCC | 2 |
| 2018 | RobinHood: Tail Latency Aware Caching - Dynamic Reallocation from Cache-Rich to Cache-Poor
Daniel S. Berger, Benjamin Berg, Timothy Zhu, Siddhartha Sen 0001, Mor Harchol-Balter |
OSDI | 3 |
| 2017 | WorkloadCompactor: reducing datacenter cost while providing tail latency SLO guaranteesabstractService providers want to reduce datacenter costs by consolidating workloads onto fewer servers. At the same time, customers have performance goals, such as meeting tail latency Service Level Objectives (SLOs). Consolidating workloads while meeting tail latency goals is challenging, especially since workloads in production environments are often bursty. To limit the congestion when consolidating workloads, customers and service providers often agree upon rate limits. Ideally, rate limits are chosen to maximize the number of workloads that can be co-located while meeting each workload's SLO. In reality, neither the service provider nor customer knows how to choose rate limits. Customers end up selecting rate limits on their own in some ad hoc fashion, and service providers are left to optimize given the chosen rate limits. Timothy Zhu, Michael A. Kozuch, Mor Harchol-Balter |
SoCC | 1 |
| 2016 | SNC-Meister: Admitting More Tenants with Tail Latency SLOsabstractMeeting tail latency Service Level Objectives (SLOs) in shared cloud networks is both important and challenging. One primary challenge is determining limits on the multi-tenancy such that SLOs are met. Doing so involves estimating latency, which is difficult, especially when tenants exhibit bursty behavior as is common in production environments. Nevertheless, recent papers in the past two years (Silo, QJump, and PriorityMeister) show techniques for calculating latency based on a branch of mathematical modeling called Deterministic Network Calculus (DNC). The DNC theory is designed for adversarial worst-case conditions, which is sometimes necessary, but is often overly conservative. Typical tenants do not require strict worst-case guarantees, but are only looking for SLOs at lower percentiles (e.g., 99th, 99.9th). Timothy Zhu, Daniel S. Berger, Mor Harchol-Balter |
SoCC | 1 |
| 2016 | TetriSched: global rescheduling with adaptive plan-ahead in dynamic heterogeneous clustersabstractTetriSched is a scheduler that works in tandem with a calendaring reservation system to continuously re-evaluate the immediate-term scheduling plan for all pending jobs (including those with reservations and best-effort jobs) on each scheduling cycle. TetriSched leverages information supplied by the reservation system about jobs' deadlines and estimated runtimes to plan ahead in deciding whether to wait for a busy preferred resource type (e.g., machine with a GPU) or fall back to less preferred placement options. Plan-ahead affords significant flexibility in handling mis-estimates in job runtimes specified at reservation time. Integrated with the main reservation system in Hadoop YARN, TetriSched is experimentally shown to achieve significantly higher SLO attainment and cluster utilization than the best-configured YARN reservation and CapacityScheduler stack deployed on a real 256 node cluster. Alexey Tumanov, Timothy Zhu, Jun Woo Park, Michael A. Kozuch, Mor Harchol-Balter, Gregory R. Ganger |
EuroSys | 2 |
| 2014 | PriorityMeister: Tail Latency QoS for Shared Networked StorageabstractMeeting service level objectives (SLOs) for tail latency is an important and challenging open problem in cloud computing infrastructures. The challenges are exacerbated by burstiness in the workloads. This paper describes PriorityMeister -- a system that employs a combination of per-workload priorities and rate limits to provide tail latency QoS for shared networked storage, even with bursty workloads. PriorityMeister automatically and proactively configures workload priorities and rate limits across multiple stages (e.g., a shared storage stage followed by a shared network stage) to meet end-to-end tail latency SLOs. In real system experiments and under production trace workloads, PriorityMeister outperforms most recent reactive request scheduling approaches, with more workloads satisfying latency SLOs at higher latency percentiles. PriorityMeister is also robust to mis-estimation of underlying storage device performance and contains the effect of misbehaving workloads. Timothy Zhu, Alexey Tumanov, Michael A. Kozuch, Mor Harchol-Balter, Gregory R. Ganger |
SoCC | 1 |
| 2013 | IOFlow: a software-defined storage architectureabstractIn data centers, the IO path to storage is long and complex. It comprises many layers or "stages" with opaque interfaces between them. This makes it hard to enforce end-to-end policies that dictate a storage IO flow's performance (e.g., guarantee a tenant's IO bandwidth) and routing (e.g., route an untrusted VM's traffic through a sanitization middlebox). These policies require IO differentiation along the flow path and global visibility at the control plane. We design IOFlow, an architecture that uses a logically centralized control plane to enable high-level flow policies. IOFlow adds a queuing abstraction at data-plane stages and exposes this to the controller. The controller can then translate policies into queuing rules at individual stages. It can also choose among multiple stages for policy enforcement. Eno Thereska, Hitesh Ballani, Greg O'Shea, Thomas Karagiannis, Antony I. T. Rowstron, Tom Talpey, Richard Black, Timothy Zhu |
SOSP | 8 |
| 2012 | SOFTScale: Stealing Opportunistically for Transient Scaling
Anshul Gandhi, Timothy Zhu, Mor Harchol-Balter, Michael A. Kozuch |
Middleware | 2 |