EDBT 2026 Demo / reviewers in the wild / expert
Moshe Hershcovitch
dblp:97/5500 · also Moshik Hershcovitch
· DBLP profile ↗
20ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-4826-4174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZipNN: Lossless Compression for AI ModelsabstractWith the growth of model sizes and the scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these. While there is a vast model compression literature deleting parts of the model weights for faster inference, we investigate a more traditional type of compression - one that represents the model in a compact form and is coupled with a decompression algorithm that returns it to its original form and size - namely lossless compression. We present ZipNN, a lossless compression tailored to neural networks. Somewhat surprisingly, we show that specific lossless compression can gain significant network and storage reduction on popular models, often saving 33% and at times reducing over 50% of the model size. We investigate the source of model compressibility and introduce specialized compression variants tailored for models that further increase the effectiveness of compression. On popular models (e.g. Llama 3) ZipNN shows space savings that are over 17% better than vanilla compression while also improving compression and decompression speeds by 62%. Using multiple workers and threads, ZipNN can achieve decompression speeds of up to 80GB/s and compression speed of up to 13GB/s. We estimate that these methods could save over an ExaByte per year of network traffic downloaded from a large model hub like Hugging Face. Moshe Hershcovitch, Andrew Wood, Leshem Choshen, Guy Girmonsky, Roy Leibovitz, Or Ozeri, Ilias Ennmouri, Michal Malka, Sang (Peter) Chin, Swaminathan Sundararaman, Danny Harnik |
CLOUD | 1 |
| 2025 | Why Paying for Storage Beats Free Networking in Cloud BurstingabstractHybrid cloud applications elastically burst to public clouds from an on-premise private cloud. In this setup only the public cloud directly charges applications for storing and processing data while the on-premise storage and network are already paid for and hence are considered free of charge. Consequently, application designers are naturally inclined to store and serve data remotely, while only paying for compute in the public cloud. Itamar Gefen, Aviad Zuck, Daniel Bransky, Moshe Hershcovitch, Danny Harnik, Dan Tsafrir |
HotStorage | 4 |
| 2025 | SkyStore: Cost-Optimized Object Storage Across Regions and CloudsabstractModern applications span multiple clouds to reduce costs, avoid vendor lock-in, and leverage low-availability resources in another cloud. However, standard object stores operate within a single cloud, forcing users to manually manage data placement across clouds, i.e., navigate their diverse APIs and handle heterogeneous costs for network and storage. This is often a complex choice: users must either pay to store objects in a remote cloud, or pay to transfer them over the network based on application access patterns and cloud provider cost offerings. To address this, we present SkyStore, a unified object store that addresses cost-optimal data management across regions and clouds. SkyStore introduces a virtual object and bucket API to hide the complexity of interacting with multiple clouds. At its core, SkyStore has a novel TTL-based data placement policy that dynamically replicates and evicts objects according to application access patterns while optimizing for lower cost. Our evaluation shows that across various workloads, SkyStore reduces the overall cost by up to 6X over academic baselines and commercial alternatives like AWS multi-region buckets. SkyStore also has comparable latency, and its availability and fault tolerance are on par with standard cloud offerings. Xiangxi Mo, Moshe Hershcovitch, Henric Zhang, Audrey Cheng, Guy Girmonsky, Gil Vernik, Michael Factor, Tiemo Bang, Soujanya Ponnapalli, Natacha Crooks, Joseph Gonzalez 0001, Danny Harnik, Ion Stoica |
Proc. VLDB Endow. | 3 |
| 2024 | Dictionary Based Cache Line CompressionabstractActive-standby mechanisms for VM high-availability demand frequent synchronization of memory and CPU state, involving the identification and transfer of "dirty" memory pages to a standby target. Building upon the granularity offered by CXL-enabled memory devices, as discussed by Waddington et al. [21], this paper proposes a dictionary-based compression method operating on 64-byte cache lines to minimize snapshot volume and synchronization latency. The method aims to transmit only necessary information required to reconstruct the memory state at the standby machine, augmented by byte grouping and cache-line partitioning techniques. We assess the compression benefits on memory access patterns across 20 benchmarks snapshots and compare our approach to standard off-the-shelf compression methods. Our findings reveal significant improvements across nearly all benchmarks, with some experiencing over a twofold enhancement compared to standard compression, while others show more moderate gains. We conduct an in-depth experimental analysis on the contribution of each method and examine the nature of the benchmarks. We ascertain that the repeating nature of cache lines across snapshots (caused by transient memory changes) and their concise representation contributes most to the size reduction, accounting for 92% of the gains. Our work paves the way for further reduction in the data transferred to standby machines, thereby enhancing VM high-availability and reducing synchronization latency. Sarel Cohen, Dalit Naor, Daniel G. Waddington, Moshe Hershcovitch |
HotStorage | 5 |
| 2023 | Fast Feature Selection with Fairness ConstraintsabstractWe study the fundamental problem of selecting optimal features for model construction. This problem is computationally challenging on large datasets, even with the use of greedy algorithm variants. To address this challenge, we extend the adaptive query model, recently proposed for the greedy forward selection for submodular functions, to the faster paradigm of Orthogonal Matching Pursuit for non-submodular functions. The proposed algorithm achieves exponentially fast parallel run time in the adaptive query model, scaling much better than prior work. Furthermore, our extension allows the use of downward-closed constraints, which can be used to encode certain fairness criteria into the feature selection process. We prove strong approximation guarantees for the algorithm based on standard assumptions. These guarantees are applicable to many parametric models, including Generalized Linear Models. Finally, we demonstrate empirically that the proposed algorithm competes favorably with state-of-the-art techniques for feature selection, on real-world and synthetic datasets. Francesco Quinzan, Rajiv Khanna, Moshe Hershcovitch, Sarel Cohen, Daniel G. Waddington, Tobias Friedrich 0001, Michael W. Mahoney |
AISTATS | 3 |
| 2023 | Cache Line Deltas CompressionabstractSynchronization of replicated data and program state is an essential aspect of application fault-tolerance. Current solutions use virtual memory mapping to identify page writes and replicate them at the destination. This approach has limitations because the granularity is restricted to a minimum of 4KiB per page, which may result in more data being replicated. Motivated by the emerging CXL hardware, we expand on the work Waddington, et al. [SoCC 22] by evaluating popular compression algorithms on VM snapshot data at cache line granularity. We measure the compression ratio vs. the compression time and present our conclusions. Sarel Cohen, Dalit Naor, Daniel G. Waddington, Moshe Hershcovitch |
SYSTOR | 5 |
| 2023 | Prefix Siphoning: Exploiting LSM-Tree Range Filters For Information Disclosure
Adi Kaufman, Moshe Hershcovitch, Adam Morrison 0001 |
USENIX ATC | 2 |
| 2023 | RLS Side Channels: Investigating Leakage of Row-Level Security Protected Data Through Query Execution TimeabstractMany modern use cases of relational databases involve multi-tenancy. To allow a tenant to only access its data, relational database systems (RDBMSs) introduced row-level security (RLS). RLS enables specifying per-row access controls, which the database enforces by rewriting tenant queries to add an RLS policy filter that filters out rows the tenant is not allowed to view. Unfortunately, while RLS blocks queries from returning unauthorized data, side-effects of query execution can form a side-channel that leaks information about such secret data. This paper investigates how RLS query execution time can leak information about rows that the querying tenant is restricted from viewing. We show that in PostgreSQL and SQL Server, an attacker can craft index-using queries to learn whether a value they are not authorized to view exists in an RLS-protected table, and in some cases, how many times such a value exists in the table. Our attack succeeds in a realistic cloud setting: we successfully attack managed PostgreSQL and SQL Server database instances on AWS from virtual machines in the same and different data centers. To block the RLS time side-channel, we design a data-oblivious query scheme for the case of unique keys. We also analyze the trade-offs created by the data-oblivious approach for non-unique keys. To facilitate the evaluation of RLS attacks and defenses, we introduce a benchmark that supports multi-tenancy and RLS, which are not supported by established benchmarks such as YCSB. We implement our solution in PostgreSQL and show that it achieves security with minimal performance impact. Chen Dar, Moshe Hershcovitch, Adam Morrison 0001 |
Proc. ACM Manag. Data | 2 |
| 2022 | A case for using cache line deltas for high frequency VM snapshottingabstractActive-standby schemes for Virtual Machine (VM) high availability require periodic synchronization of memory and CPU state. The most common approach to synchronization is to use page tables and software to identify "dirty" memory pages at the source and in turn copy them to the target via a network or interconnect. However, this approach results in significanct page table traversal and data copying overhead, resulting in considerable VM downtime. A principal contributor to this overhead is that many applications using this approach incur data copy-amplification as a result of copying more data than is necessary; this arises because of the processor's virtual memory system design in which memory pages are 4KiB or larger. Daniel G. Waddington, Moshe Hershcovitch, Swaminathan Sundararaman, Clem Dickey |
SoCC | 2 |
| 2022 | Elastic Indexes: Dynamic Space vs. Query Efficiency Tuning for In-Memory Database Indexing
Moshe Hershcovitch, Artem Khyzha, Daniel G. Waddington, Adam Morrison 0001 |
EDBT | 1 |
| 2022 | Evaluating compressed indexes in DBMSabstractIn-memory database management systems (DBMSs) are an essential part of real-world applications. They store their entire data in memory, and thus their performance is much higher than standard DBMS that uses slow block-based storage. The memory footprint is the essential resource in such systems, while the database indexes consume a large portion of the memory and can reach up to 50% of total memory consumption [2]. Oz Anani, Gal Lushi, Moshe Hershcovitch, Adam Morrison 0001 |
SYSTOR | 3 |
| 2022 | System-level crash safe sorting on persistent memoryabstractSorting is a fundamental operation in software systems. An example for that is a prepossessing phase before executing analytics operations. Omri Arad, Yoav Ben Shimon, Ron Zadicario, Daniel G. Waddington, Moshe Hershcovitch, Adam Morrison 0001 |
SYSTOR | 5 |
| 2021 | PyMM: Heterogeneous Memory Programming for Python Data ScienceabstractWhile persistent memory (PMEM) is a promising technology, leveraging it with legacy applications is non-trivial. This is primarily because legacy applications assume all memory is volatile and there is no notion of crash-consistency or state recovery. As new types of persistent and intelligent memory emerge, propelled by the CXL standard, the problem of integration and adoption remains. Daniel G. Waddington, Moshe Hershcovitch, Clem Dickey |
PLOS@SOSP | 2 |
| 2021 | DeCorus-NSA: detection and correlation of unusual signals for network syslog analyticsabstractThe management of large data centre (DC) network infrastructure confronts Network Reliability Engineers (NRE) with challenges. A single DC at a modern cloud services provider can host thousands of network devices. The syslog messages generated by these devices are an important type of monitoring data to detect and diagnose failures. Devices in a single DC produce millions of syslog messages per day in a variety of formats. David Ohana, Bruno Wassermann, Moshe Hershcovitch, Elliot K. Kolodner, Michal Malka, Eran Raichstein, Ronen Schaffer, Robert Shahla |
SYSTOR | 3 |
| 2020 | Sketching Volume Capacities in Deduplicated StorageabstractThe adoption of deduplication in storage systems has introduced significant new challenges for storage management. Specifically, the physical capacities associated with volumes are no longer readily available. In this work, we introduce a new approach to analyzing capacities in deduplicated storage environments. We provide sketch-based estimations of fundamental capacity measures required for managing a storage system: How much physical space would be reclaimed if a volume or group of volumes were to be removed from a system (the reclaimable capacity) and how much of the physical space should be attributed to each of the volumes in the system (the attributed capacity). Our methods also support capacity queries for volume groups across multiple storage systems, e.g., how much capacity would a volume group consume after being migrated to another storage system? We provide analytical accuracy guarantees for our estimations as well as empirical evaluations. Our technology is integrated into a prominent all-flash storage array and exhibits high performance even for very large systems. We also demonstrate how this method opens the door for performing placement decisions at the data-center level and obtaining insights on deduplication in the field. Danny Harnik, Moshe Hershcovitch, Yosef Shatsky, Amir Epstein, Ronen I. Kat |
ACM Trans. Storage | 2 |
| 2019 | Sketching Volume Capacities in Deduplicated Storage
Danny Harnik, Moshe Hershcovitch, Yosef Shatsky, Amir Epstein, Ronen I. Kat |
FAST | 2 |
| 2018 | PM aware storage engine for MongoDBabstractWith the maturity of Persistant Memories (PM) such as storage class memory technologies, e.g., STT-MRAM, PCM, ReRAM and 3DXpoint, we expect to see practical implementation of data structures, data stores and databases for use-cases such as IoT, mobile, and cloud. We developed a PM-aware storage engine for MongoDB which leverages PM hardware capabilities such as byte addressability and persistency. With our storage engine we see improved latency, less write amplification, less capacity and simpler implementation due to the fact that some code paths become unnecessary compared to past implementations. Moshe Hershcovitch, Revital Eres, Adam J. McPadden |
SYSTOR | 1 |
| 2015 | Minimal indices for predecessor search
Sarel Cohen, Amos Fiat, Moshe Hershcovitch, Haim Kaplan |
Inf. Comput. | 3 |
| 2013 | Minimal Indices for Successor Search - (Extended Abstract)
Sarel Cohen, Amos Fiat, Moshe Hershcovitch, Haim Kaplan |
MFCS | 3 |
| 2013 | I/O Efficient Dynamic Data Structures for Longest Prefix Queries
Moshe Hershcovitch, Haim Kaplan |
Algorithmica | 1 |