EDBT 2026 Demo / reviewers in the wild / expert
Deepavali Bhagwat
dblp:60/4735
· DBLP profile ↗
9ranked-venue papers in the field
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 4 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | CNSBench: A Cloud Native Storage Benchmark
Alex Merenstein, Vasily Tarasov, Ali Anwar 0001, Deepavali Bhagwat, Julie Lee, Lukas Rupprecht, Dimitrios Skourtis, Erez Zadok |
FAST | 4 |
| 2020 | Improving Reproducibility of Data Science Pipelines through Transparent Provenance CaptureabstractData science has become prevalent in a large variety of domains. Inherent in its practice is an exploratory, probing, and fact finding journey, which consists of the assembly, adaptation, and execution of complex data science pipelines. The trustworthiness of the results of such pipelines rests entirely on their ability to be reproduced with fidelity, which is difficult if pipelines are not documented or recorded minutely and consistently. This difficulty has led to a reproducibility crisis and presents a major obstacle to the safe adoption of the pipeline results in production environments. The crisis can be resolved if the provenance for each data science pipeline is captured transparently as pipelines are executed. However, due to the complexity of modern data science pipelines, transparently capturing sufficient provenance to allow for reproducibility is challenging. As a result, most existing systems require users to augment their code or use specific tools to capture provenance, which hinders productivity and results in a lack of adoption. In this paper, we present Ursprung, 1 a transparent provenance collection system designed for data science environments. 2 The Ursprung philosophy is to capture provenance and build lineage by integrating with the execution environment to automatically track static and runtime configuration parameters of data science pipelines. Rather than requiring data scientists to make changes to their code, Ursprung records basic provenance information from system-level sources and combines it with provenance from application-level sources (e.g., log files, stdout), which can be accessed and recorded through a domain-specific language. In our evaluation, we show that Ursprung is able to capture sufficient provenance for a variety of use cases and only adds an overhead of up to 4%. Lukas Rupprecht, James C. Davis 0001, Constantine Arnold, Yaniv Gur, Deepavali Bhagwat |
Proc. VLDB Endow. | 5 |
| 2019 | Ursprung: Provenance for Large-Scale Analytics EnvironmentsabstractModern analytics has produced wonders, but reproducing and verifying these wonders is difficult. Data provenance helps to solve this problem by collecting information on how data is created and accessed. Although provenance collection techniques have been used successfully on a smaller scale, tracking provenance in large-scale analytics environments is challenging due to the scale of provenance generated and the heterogeneous domains. Without provenance, analysts struggle to keep track of and reproduce their analyses. We demonstrate Ursprung, a provenance collection system specifically targeted at such environments. Ursprung transparently collects the minimal set of system-level provenance required to track the relationships between data and processes. To collect domain specific provenance, Usprung enables users to specify capture rules to curate application-specific logs, intermediate results etc. To reduce storage overhead and accelerate queries, it uses event hierarchies to synthesize raw provenance into compact summaries. Lukas Rupprecht, James C. Davis 0001, Constantine Arnold, Alexander L. R. Lubbock, Darren R. Tyson, Deepavali Bhagwat |
SIGMOD Conference | 6 |
| 2015 | A Practical Implementation of Clustered Fault Tolerant Write Acceleration in a Virtualized Environment
Deepavali Bhagwat, Mahesh Patil, Michal Ostrowski, Murali Vilayannur, Woon Jung, Chethan Kumar |
FAST | 1 |
| 2013 | Improving restore speed for backup systems that use inline chunk-based deduplication
Mark Lillibridge, Kave Eshghi, Deepavali Bhagwat |
FAST | 3 |
| 2009 | Sparse Indexing: Large Scale, Inline Deduplication Using Sampling and Locality
Mark Lillibridge, Kave Eshghi, Deepavali Bhagwat, Vinay Deolalikar, Greg Trezis, Peter Camble |
FAST | 3 |
| 2007 | Content-based document routing and index partitioning for scalable similarity-based searches in a large corpusabstractWe present a document routing and index partitioning scheme for scalable similarity-based search of documents in a large corpus. We consider the case when similarity-based search is performed by finding documents that have features in common with the query document. While it is possible to store all the features of all the documents in one index, this suffers from obvious scalability problems. Our approach is to partition the feature index into multiple smaller partitions that can be hosted on separate servers, enabling scalable and parallel search execution. When a document is ingested into the repository, a small number of partitions are chosen to store the features of the document. To perform similarity-based search, also, only a small number of partitions are queried. Our approach is stateless and incremental. The decision as to which partitions the features of the document should be routed to (for storing at ingestion time, and for similarity based search at query time) is solely based on the features of the document. Deepavali Bhagwat, Kave Eshghi, Pankaj Mehra |
KDD | 1 |
| 2005 | An annotation management system for relational databases
Deepavali Bhagwat, Laura Chiticariu, Wang Chiew Tan, Gaurav Vijayvargiya |
VLDB J. | 1 |
| 2004 | An Annotation Management System for Relational Databases
Deepavali Bhagwat, Laura Chiticariu, Wang Chiew Tan, Gaurav Vijayvargiya |
VLDB | 1 |