VLDB 2026 Research / reviewers in the wild / expert
Guy Khazma
dblp:241/5858
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0008-1364-5162ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Asymmetric Linearizable Local ReadsabstractMany linearizable local read algorithms have been proposed to minimize the read latency of strongly consistent distributed databases deployed in geo-distributed networks. These algorithms do so by enabling reads to be performed immediately against any process' copy of the database in the best case. However, as our analysis shows, worst-case read latency at every process with all existing algorithms is at least the network's relative diameter in terms of the maximum message delay minus a known lower bound on message delay between any two processes. We then show that by leveraging the asymmetric message delays of geo-distributed networks, worst-case read latency can be below the network's relative diameter at processes close to the leader or the network's center by presenting two new linearizable local read algorithms. Our experimental evaluation shows that these new algorithms reduce worst-case read latency by up to 50x compared to existing ones. Myles Thiessen, Guy Khazma, Sam Toueg, Eyal de Lara |
Proc. VLDB Endow. | 2 |
| 2023 | Refactoring ETL Flows in The WildabstractIn modern data-driven ecosystems, Extract, Transform, Load (ETL) flows serve as the backbone of data integration pipelines. These flows facilitate the seamless movement of data across disparate systems and formats, streamlining processes that range from data acquisition to preparation for analysis. However, the pervasive use of ETL flows introduces a pressing challenge-how to bound the maintenance cost of an ever-expanding number of flows. In this paper, we describe an end-to-end prototype for ETL flow refactoring, aimed at reducing the maintenance cost, which keeps the human in the loop for refactoring decisions. Our prototype adopts and significantly extends the gSpan Frequent Subgraph Mining (FSM) algorithm to apply it to real-world ETL use cases in the context of the IBM DataStage™ data integration tool. We report on real customer workloads, share their statistics and evaluate the performance of our prototype. We found potential for up to 32% maintenance cost reduction on the use cases we analyzed after removing duplicate flows. We also share an anonymized version of the workloads with the research community. Dolev Adas, Ohad Eytan, Guy Khazma, Josep Sampé, Paula Ta-Shma |
IEEE Big Data | 3 |
| 2022 | DSON: JSON CRDT Using Delta-Mutations For Document StoresabstractWe propose DSON, a space efficient δ-based CRDT approach for distributed JSON document stores, enabling high availability at a global scale, while providing strong eventual consistency guarantees. We define the semantics of our CRDT based approach formally, and prove its correctness and convergence. Previous approaches optimize for collaborative document editing and store metadata proportional to the number of updates to a document, which is not acceptable for long lived document management. The metadata stored with our approach is bounded by O ( k 2 D + n log n ), where n is the number of replicas, D is the number of document elements, and k ≤ n is the number of concurrent document updates. We also implement our approach[37] and demonstrate its space efficiency empirically. Experimental analysis shows that the metadata stored is typically significantly less than the worst case. This provides the basis for robust highly available distributed document stores with well defined semantics and safety guarantees, relieving application developers from the burden of conflict resolution. Arik Rinberg, Tomer Solomon, Roee Shlomo, Guy Khazma, Gal Lushi, Idit Keidar, Paula Ta-Shma |
Proc. VLDB Endow. | 4 |
| 2020 | Extensible Data SkippingabstractData skipping reduces I/O for SQL queries by skipping over irrelevant data objects (files) based on their metadata. We extend this notion by allowing developers to define their own data s kipping metadata types and indexes using a flexible A PI. Our framework i s t he first to natively support data skipping for arbitrary data types (e.g. geospatial, logs) and queries with User Defined Functions ( UDFs). We integrated our framework with Apache Spark and it is now deployed across multiple products/services at IBM. We present our extensible data skipping APIs, discuss index design, and implement various metadata indexes, requiring only around 30 lines of additional code per index. In particular we implement data skipping for a third party library with geospatial UDFs and demonstrate speedups of two orders of magnitude. Our centralized metadata approach provides a x3.6 speed up even when compared to queries which are rewritten to exploit Parquet min/max metadata. We demonstrate that extensible data skipping is applicable to broad class of applications, where user defined indexes achieve significant speedups and cost savings with very low development cost. Paula Ta-Shma, Guy Khazma, Gal Lushi, Oshrit Feder |
IEEE BigData | 2 |
| 2019 | Big data skipping in the cloudabstractAccording to today's best practices, cloud compute and storage services should be deployed and managed independently. However, this generates a problem for big data analytics in the cloud: potentially huge datasets need to be shipped from the storage service to the compute service to analyse the data. To address this, minimizing the amount of data sent across the network is critical to achieve good performance and low cost. Data skipping is a technique which achieves this for SQL style analytics on structured data. Oshrit Feder, Guy Khazma, Gal Lushi, Yosef Moatti, Paula Ta-Shma |
SYSTOR | 2 |