VLDB 2026 Research / reviewers in the wild / expert
Harshal A. Chaudhari
dblp:213/7820
· DBLP profile ↗
6ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-0444-4915ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AXE: A Task Decomposition Approach to Learned LSM Tuning
Andy Huynh, Anwesha Saha, Harshal A. Chaudhari, Manos Athanassoulis |
Proc. VLDB Endow. | 3 |
| 2024 | Towards flexibility and robustness of LSM treesabstractAbstract Log-structured merge trees (LSM trees) are increasingly used as part of the storage engine behind several data systems, and are frequently deployed in the cloud. As the number of applications relying on LSM-based storage backends increases, the problem of performance tuning of LSM trees receives increasing attention. We consider both nominal tunings—where workload and execution environment are accurately known a priori—and robust tunings—which consider uncertainty in the workload knowledge. This type of workload uncertainty is common in modern applications, notably in shared infrastructure environments like the public cloud. To address this problem, we introduce Endure , a new paradigm for tuning LSM trees in the presence of workload uncertainty. Specifically, we focus on the impact of the choice of compaction policy, size ratio, and memory allocation on the overall performance. Endure considers a robust formulation of the throughput maximization problem and recommends a tuning that offers near-optimal throughput when the executed workload is not the same, instead in a neighborhood of the expected workload. Additionally, we explore the robustness of flexible LSM designs by proposing a new unified design called K-LSM that encompasses existing designs. We deploy our robust tuning system, Endure , on a state-of-the-art key-value store, RocksDB, and demonstrate throughput improvements of up to 5 $$\times $$ × in the presence of uncertainty. Our results indicate that the tunings obtained by Endure are more robust than tunings obtained under our expanded LSM design space. This indicates that robustness may not be inherent to a design, instead, it is an outcome of a tuning process that explicitly accounts for uncertainty. Andy Huynh, Harshal A. Chaudhari, Evimaria Terzi, Manos Athanassoulis |
VLDB J. | 2 |
| 2022 | Endure: A Robust Tuning Paradigm for LSM Trees Under Workload UncertaintyabstractLog-Structured Merge trees (LSM trees) are increasingly used as the storage engines behind several data systems, frequently deployed in the cloud. Similar to other database architectures, LSM trees consider information about the expected workload (e.g., reads vs. writes, point vs. range queries) to optimize their performance via tuning. However, operating in a shared infrastructure like the cloud comes with workload uncertainty due to the fast-evolving nature of modern applications. Systems with static tuning discount the variability of such hybrid workloads and hence provide an inconsistent and overall suboptimal performance. To address this problem, we introduce Endure - a new paradigm for tuning LSM trees in the presence of workload uncertainty. Specifically, we focus on the impact of the choice of compaction policies, size ratio, and memory allocation on the overall performance. Endure considers a robust formulation of the throughput maximization problem and recommends a tuning that maximizes the worst-case throughput over the neighborhood of each expected workload. Additionally, an uncertainty tuning parameter controls the size of this neighborhood, thereby allowing the output tunings to be conservative or optimistic. Through both model-based and extensive experimental evaluations of Endure in the state-of-the-art LSM-based storage engine, RocksDB, we show that the robust tuning methodology consistently outperforms classical tuning strategies. The robust tunings output by Endure lead up to a 5X improvement in throughput in the presence of uncertainty. On the flip side, Endure tunings have negligible performance loss when the observed workload exactly matches the expected one. Andy Huynh, Harshal A. Chaudhari, Evimaria Terzi, Manos Athanassoulis |
Proc. VLDB Endow. | 2 |
| 2020 | Learn to Earn: Enabling Coordination Within a Ride-Hailing FleetabstractThe problem of optimizing social welfare objectives on multi-sided ride-hailing platforms such as Uber, Lyft, etc., is challenging, due to misalignment of objectives between drivers, passengers, and the platform itself. An ideal solution aims to minimize the response time for each hyperlocal passenger ride request, while simultaneously maintaining high demand satisfaction and supply utilization across the entire city. Economists tend to rely on dynamic pricing mechanisms that stifle price-sensitive excess demand and resolve supply-demand imbalances that emerge in specific neighborhoods. In contrast, computer scientists primarily view it as a demand prediction problem with the goal of preemptively repositioning supply to such neighborhoods using black-box coordinated multi-agent deep reinforcement learning-based approaches. Here, we introduce explainability in the existing supply-repositioning approaches by establishing the need for coordination between the drivers at specific locations and times. Explicit need-based coordination allows our framework to use a simpler non-deep reinforcement learning-based approach, thereby enabling it to explain its recommendations ex-post. Moreover, it provides envy-free recommendations i.e., drivers at the same location and time do not envy one another's expected future earnings. Our experimental evaluation demonstrates the effectiveness, robustness, and generalizability of our framework. Finally, in contrast to previous works, we make available a reinforcement learning environment for end-to-end reproducibility of our work and to encourage future comparative studies. Harshal A. Chaudhari, John W. Byers, Evimaria Terzi |
IEEE BigData | 1 |
| 2018 | Markov Chain MonitoringabstractIn networking applications, one often wishes to obtain estimates about the number of objects at different parts of the network (e.g., the number of cars at an intersection of a road network or the number of packets expected to reach a node in a computer network) by monitoring the traffic in a small number of network nodes or edges. We formalize this task by defining the Markov Chain Monitoring problem. Given an initial distribution of items over the nodes of a Markov chain, we wish to estimate the distribution of items at subsequent times. We do this by asking a limited number of queries that retrieve, for example, how many items transitioned to a specific node or over a specific edge at a particular time. We consider different types of queries, each defining a different variant of the Markov Chain Monitoring. For each variant, we design efficient algorithms for choosing the queries that make our estimates as accurate as possible. In our experiments with synthetic and real datasets, we demonstrate the efficiency and the efficacy of our algorithms in a variety of settings. Harshal A. Chaudhari, Michael Mathioudakis, Evimaria Terzi |
SDM | 1 |
| 2018 | Putting Data in the Driver's Seat: Optimizing Earnings for On-Demand Ride-HailingabstractOn-demand ride-hailing platforms like Uber and Lyft are helping reshape urban transportation, by enabling car owners to become drivers for hire with minimal overhead. Although there are many studies that consider ride-hailing platforms holistically, e.g., from the perspective of supply and demand equilibria, little emphasis has been placed on optimization for the individual, self-interested drivers that currently comprise these fleets. While some individuals drive opportunistically either as their schedule allows or on a fixed schedule, we show that strategic behavior regarding when and where to drive can substantially increase driver income. In this paper, we formalize the problem of devising a driver strategy to maximize expected earnings, describe a series of dynamic programming algorithms to solve these problems under different sets of modeled actions available to the drivers, and exemplify the models and methods on a large scale simulation of driving for Uber in NYC. In our experiments, we use a newly-collected dataset that combines the NYC taxi rides dataset along with Uber API data, to build time-varying traffic and payout matrices for a representative six-month time period in greater NYC. From this input, we can reason about prospective itineraries and payoffs. Moreover, the framework enables us to rigorously reason about and analyze the sensitivity of our results to perturbations in the input data. Among our main findings is that repositioning throughout the day is key to maximizing driver earnings, whereas »chasing surge' is typically misguided and sometimes a costly move. Harshal A. Chaudhari, John W. Byers, Evimaria Terzi |
WSDM | 1 |