VLDB 2026 Research / reviewers in the wild / expert
Grisha Weintraub
dblp:201/4808 · also Gregory Weintraub
· DBLP profile ↗
10ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0003-4823-4757ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 first-author · 7 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Cloud Data Lake Queries by Minimizing the Query Coverage SetabstractCloud data lakes provide a modern solution for managing large volumes of data. The fundamental principle behind these systems is the separation of compute and storage layers. In this architecture, inexpensive cloud storage is utilized for data storage, while compute engines are employed to perform analytics on this data in an “on-demand” mode. However, to execute any calculations on the data, it must be transferred from the storage layer to the compute layer over the network for each query. This transfer can negatively impact calculation performance and requires significant network bandwidth. In our work, we examine various strategies to enhance query performance within a cloud data lake architecture. We begin by formalizing the problem and proposing a straightforward yet robust theoretical framework that clearly outlines the associated trade-offs. Central to our framework is the concept of a “query coverage set,” which is defined as the collection of files that need to be accessed from storage to fulfill a specific query. Our objective is to identify the minimal coverage set for each query and execute the query exclusively on this subset of files. This approach enables us to significantly improve query performance across three different domains: indexing, caching, and genetic data. Grisha Weintraub, Ehud Gudes, Shlomi Dolev |
ICDE | 1 |
| 2025 | From Alert Fatigue to Trusted Insights: A Simple Non-ML Anomaly Detection ExperienceabstractDistinguishing actionable anomalies from noise in monitoring metrics is a critical operational challenge. In our experience, static thresholds often resulted in overwhelming alert fatigue, while pilot projects utilizing ML-based models proved to be uninterpretable. We implemented an interpretable system based on simple rules co-designed with our support team, focusing on sustained, meaningful deviations from a seasonal baseline. The result was a dramatic shift in operator trust, with feedback indicating that "all alerts are good". Avi Illouz, Leonid Rise, Michal Lazarovitz, Eli Shemesh, Grisha Weintraub |
SYSTOR | 5 |
| 2024 | Coverage-Based Caching in Cloud Data LakesabstractCloud data lakes are a modern approach to handling large volumes of data. They separate the compute and storage layers, making them highly scalable and cost-effective. However, query performance in cloud data lakes could be faster, and various efforts have been made to enhance it in recent years. We introduce our approach to this problem, which is based on a novel caching technique where instead of caching actual data, we cache metadata called a coverage set. Grisha Weintraub, Ehud Gudes, Shlomi Dolev |
SYSTOR | 1 |
| 2024 | Integrity Verification in Cloud Data LakesabstractCloud data lakes support storage and querying at scale. However, traditional data integrity methods do not apply to them due to a different system model. We propose a novel completeness verification protocol based on a data lake partitioning scheme. Grisha Weintraub, Leonid Rise, Eli Shemesh, Avraham Illouz |
SYSTOR | 1 |
| 2024 | Optimizing Cloud Data Lake Queries With a Balanced Coverage PlanabstractCloud data lakes emerge as an inexpensive solution for storing very large amounts of data. The main idea is the separation of compute and storage layers. Thus, cheap cloud storage is used for storing the data, while compute engines are used for running analytics on this data in “on-demand” mode. However, to perform any computation on the data in this architecture, the data should be moved from the storage layer to the compute layer over the network for each calculation. Obviously, that hurts calculation performance and requires huge network bandwidth. In this paper, we study different approaches to improve query performance in a data lake architecture. We define an optimization problem that can provably speed up data lake queries. We prove that the problem is NP-hard and suggest heuristic approaches. Then, we demonstrate through the experiments that our approach is feasible and efficient (up to ×30 query execution time improvement based on the TPC-H benchmark). Grisha Weintraub, Ehud Gudes, Shlomi Dolev, Jeffrey D. Ullman |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Analyzing large-scale genomic data with cloud data lakesabstractIn recent years there is huge influx of genomic data and a growing need for its analysis, yet existing genomic databases do not allow easy accessibility. We developed a pipeline that continuously pre-processes raw human genetic data. The data is then stored in a cloud data lake and can be accessed via a simple and intuitive web service and API. Grisha Weintraub, Noam Hadar, Ehud Gudes, Shlomi Dolev, Ohad S. Birk |
SYSTOR | 1 |
| 2022 | Integrity verification in cloud key-value storesabstractDatabase-as-a-Service (DBaaS) is a common approach for storing data in the cloud. However, this approach introduces several concerns regarding data integrity. As a client does not maintain the data, its completeness and correctness might be compromised by the cloud provider or a malicious entity that penetrated the cloud. We are introducing a novel method for verifying data integrity in cloud databases with a focus on key-value stores. Grisha Weintraub, Leonid Rise, Alon Kadosh |
SYSTOR | 1 |
| 2021 | Indexing cloud data lakes within the lakesabstractCloud data lakes are a modern approach for storing large amounts of data in a convenient and inexpensive way. The main idea is the separation of compute and storage layers. However, to perform analytics on the data in this architecture, the data should be moved from the storage layer to the compute layer over the network for each calculation. Obviously, that hurts calculation performance and requires huge network bandwidth. We are exploring different approaches for adding indexing to the cloud data lakes with the goal of reducing the amounts of data read from the storage, and as a result, improving query execution time. Grisha Weintraub, Ehud Gudes, Shlomi Dolev |
SYSTOR | 1 |
| 2018 | Data Integrity Verification in Column-Oriented NoSQL Databases
Grisha Weintraub, Ehud Gudes |
DBSec | 1 |
| 2017 | Crowdsourced Data Integrity Verification for Key-Value Stores in the CloudabstractThanks to their high availability, scalability, and usability, cloud databases have become one of the dominant cloud services. However, since cloud users do not physically possess their data, data integrity may be at risk. In this paper, we present a novel protocol that utilizes crowdsourcing paradigm to provide practical data integrity assurance in key-value cloud databases. The main advantage of our protocol over previous work is its high applicability - as opposed to existing approaches, our scheme does not require any system changes on the cloud side and thus can be applied directly to any existing system. We demonstrate the feasibility of our scheme by a prototype implementation and its evaluation. Grisha Weintraub, Ehud Gudes |
CCGrid | 1 |