VLDB 2026 Research / reviewers in the wild / expert
Shai Bergman
dblp:184/8323 · also Shai Aviram Bergman
· DBLP profile ↗
13ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-3484-1452ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read MappingabstractGenome sequencing has become a central focus in computational biology due to its critical role in applications such as personalized medicine, disease outbreak tracking, and evolutionary research. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments known as reads. Extracting meaningful biological insights from these reads requires a computationally intensive step called read mapping, where each read is aligned to a reference genome. Read mapping for short reads comes in two forms: single-end and paired-end, with the latter being more prevalent due to its higher accuracy and support for advanced analysis. Read mapping remains a major performance bottleneck in genome analysis as a result of the extensive use of computationally intensive dynamic programming. Prior efforts have attempted to mitigate this cost by employing filters to identify and potentially discard computationally expensive matches and leveraging hardware accelerators to speed up the computations. While partially effective, these approaches have limitations. In particular, existing filters are often ineffective for paired-end reads, as they evaluate each read independently and exhibit relatively low filtering ratios. In this work, we propose GenPairX, a hardware-algorithm codesigned accelerator that efficiently minimizes the computational load of paired-end read mapping while enhancing the throughput of memory-intensive operations. GenPairX introduces: (1) a novel filtering algorithm that jointly considers both reads in a pair to improve filtering effectiveness, and a lightweight alignment algorithm to replace most of the computationally expensive dynamic programming operations, and (2) two specialized hardware mechanisms to support the proposed algorithms. The proposed hardware addresses the high memory bandwidth demands of the read filtering process via orchestration of memory accesses over high-bandwidth memory channels, and accelerates the alignment of candidate reads via simple vectorized logical XOR operators. Our evaluations show that GenPairX delivers substantial performance improvements over state-of-the-art solutions, achieving$1575 \times$and$1.43 \times$higher throughput per watt compared to leading CPU-based and accelerator-based read mappers, respectively, all without compromising accuracy. Julien Eudine, Renzo Andri, Can Firtina, Mohammad Sadrosadati, Nika Mansouri-Ghiasi, Konstantina Koliogeorgi, Anirban Nag, Arash Tavakkol, Haiyu Mao, Onur Mutlu, Shai Bergman, Ji Zhang 0035 |
HPCA | 13 |
| 2025 | Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
Shai Bergman, Anne-Marie Kermarrec, Diana Petrescu, Rafael Pires 0001, Mathis Randl, Martijn de Vos, Ji Zhang 0035 |
Middleware | 1 |
| 2025 | Accelerating Nested Virtualization with HyperTurtle
Ori Ben Zur, Jakob Krebs, Shai Bergman, Mark Silberstein |
USENIX ATC | 3 |
| 2024 | Composable Storage Servers: A Storage Paradigm for Disaggregated SystemsabstractDisaggregated and composable data centers optimize resource allocation and mitigate overprovisioning through dynamic resource management. While significant research has concentrated on disaggregating components and composing compute servers, the application of these principles to storage servers within disaggregated data centers remains underexplored. Traditional storage servers are often overprovisioned to accommodate a wide range of scenarios and workloads, resulting in designs that conflict with the principles of composable data centers, which prioritize efficiency by minimizing resource over-provisioning across components. This paper presents the concept of Composable Storage Servers, a design approach for storage solutions within disaggregated data centers. By applying the principles of resource disaggregation, this approach mitigates the overprovisioning of resources in storage servers. Central to this vision is the Core Storage Node, which integrates essential storage functionalities into a unified component, consistent with composable infrastructure principles. Our prototype and real-world deployment of the core storage node show its effectiveness at substituting local SSDs, achieving a 30% reduction in storage costs. Shai Bergman, Onur Mutlu, Wu Yong, Keji Huang, Ji Zhang 0035 |
NAS | 1 |
| 2023 | Reducing The Virtual Memory Overhead in Nested VirtualizationabstractVirtualization has become a critical aspect of modern computing, and with the advent of virtualization-based containers, fast nested virtualization has become increasingly important. Nested virtualization is implemented by emulating virtualization capabilities to the guest host which can result in significant overhead. Another source of overheads in virtualization stems from the address translation mechanisms employed to implement virtualization, which usually causes a mix of slower address translation, frequently trapping guests, and loss of granularity in page tables. Our research focuses on using guest-managed physical memory with the use of per-VM memory tags for checking each VMs' access permissions. Ori Ben Zur, Shai Bergman, Mark Silberstein |
SYSTOR | 2 |
| 2023 | Translation Pass-Through for Near-Native Paging Performance in VMs
Shai Bergman, Mark Silberstein, Takahiro Shinagawa, Peter R. Pietzuch, Lluís Vilanova |
USENIX ATC | 1 |
| 2023 | ZNSwap: un-Block your SwapabstractWe introduce ZNSwap , a novel swap subsystem optimized for the recent Zoned Namespace (ZNS) SSDs. ZNSwap leverages ZNS’s explicit control over data management on the drive and introduces a space-efficient host-side Garbage Collector (GC) for swap storage co-designed with the OS swap logic. ZNSwap enables cross-layer optimizations, such as direct access to the in-kernel swap usage statistics by the GC to enable fine-grain swap storage management, and correct accounting of the GC bandwidth usage in the OS resource isolation mechanisms to improve performance isolation in multi-tenant environments. We evaluate ZNSwap using standard Linux swap benchmarks and two production key-value stores. ZNSwap shows significant performance benefits over the Linux swap on traditional SSDs, such as stable throughput for different memory access patterns, and 10× lower 99th percentile latency and 5× higher throughput for memcached key-value store under realistic usage scenarios. Shai Bergman, Niklas Cassel, Matias Bjørling, Mark Silberstein |
ACM Trans. Storage | 1 |
| 2022 | Slashing the disaggregation tax in heterogeneous data centers with FractOSabstractDisaggregated heterogeneous data centers promise higher efficiency, lower total costs of ownership, and more flexibility for data-center operators. However, current software stacks can levy a high tax on application performance. Applications and OSes are designed for systems where local PCIe-connected devices are centrally managed by CPUs, but this centralization introduces unnecessary messages through the shared data-center network in a disaggregated system. Lluís Vilanova, Lina Maudlej, Shai Bergman, Till Miemietz, Matthias Hille, Nils Asmussen, Michael Roitzsch, Hermann Härtig, Mark Silberstein |
EuroSys | 3 |
| 2022 | Reconsidering OS memory optimizations in the presence of disaggregated memoryabstractTiered memory systems introduce an additional memory level with higher-than-local-DRAM access latency and require sophisticated memory management mechanisms to achieve cost-efficiency and high performance. Recent works focus on byte-addressable tiered memory architectures which offer better performance than pure swap-based systems. We observe that adding disaggregation to a byte-addressable tiered memory architecture requires important design changes that deviate from the common techniques that target lower-latency non-volatile memory systems. Our comprehensive analysis of real workloads shows that the high access latency to disaggregated memory undermines the utility of well-established memory management optimizations Based on these insights, we develop HotBox – a disaggregated memory management subsystem for Linux that strives to maximize the local memory hit rate with low memory management overhead. HotBox introduces only minor changes to the Linux kernel while outperforming state-of-the-art systems on memory-intensive benchmarks by up to 2.25×. Shai Bergman, Priyank Faldu, Boris Grot, Lluís Vilanova, Mark Silberstein |
ISMM | 1 |
| 2022 | ZNSwap: un-Block your Swap
Shai Bergman, Niklas Cassel, Matias Bjørling, Mark Silberstein |
USENIX ATC | 1 |
| 2018 | SPIN: Seamless Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUsabstractRecent GPUs enable Peer-to-Peer Direct Memory Access ( p 2 p ) from fast peripheral devices like NVMe SSDs to exclude the CPU from the data path between them for efficiency. Unfortunately, using p 2 p to access files is challenging because of the subtleties of low-level non-standard interfaces, which bypass the OS file I/O layers and may hurt system performance. Developers must possess intimate knowledge of low-level interfaces to manually handle the subtleties of data consistency and misaligned accesses. We present SPIN , which integrates p 2 p into the standard OS file I/O stack, dynamically activating p 2 p where appropriate, transparently to the user. It combines p 2 p with page cache accesses, re-enables read-ahead for sequential reads, all while maintaining standard POSIX FS consistency, portability across GPUs and SSDs, and compatibility with virtual block devices such as software RAID. We evaluate SPIN on NVIDIA and AMD GPUs using standard file I/O benchmarks, application traces, and end-to-end experiments. SPIN achieves significant performance speedups across a wide range of workloads, exceeding p 2 p throughput by up to an order of magnitude. It also boosts the performance of an aerial imagery rendering application by 2.6× by dynamically adapting to its input-dependent file access pattern, enables 3.3× higher throughput for a GPU-accelerated log server, and enables 29% faster execution for the highly optimized GPU-accelerated image collage with only 30 changed lines of code. Shai Bergman, Tanya Brokhman, Tzachi Cohen, Mark Silberstein |
ACM Trans. Comput. Syst. | 1 |
| 2017 | SPIN: Seamless Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs
Shai Bergman, Tanya Brokhman, Tzachi Cohen, Mark Silberstein |
USENIX ATC | 1 |
| 2016 | ActivePointers: A Case for Software Address Translation on GPUsabstractModern discrete GPUs have been the processors of choice for accelerating compute-intensive applications, but using them in large-scale data processing is extremely challenging. Unfortunately, they do not provide important I/O abstractions long established in the CPU context, such as memory mapped files, which shield programmers from the complexity of buffer and I/O device management. However, implementing these abstractions on GPUs poses a problem: the limited GPU virtual memory system provides no address space management and page fault handling mechanisms to GPU developers, and does not allow modifications to memory mappings for running GPU programs. We implement ActivePointers, a software address translation layer and paging system that introduces native support for page faults and virtual address space management to GPU programs, and enables the implementation of fully functional memory mapped files on commodity GPUs. Files mapped into GPU memory are accessed using active pointers, which behave like regular pointers but access the GPU page cache under the hood, and trigger page faults which are handled on the GPU. We design and evaluate a number of novel mechanisms, including a translation cache in hardware registers and translation aggregation for deadlock-free page fault handling of threads in a single warp. We extensively evaluate ActivePointers on commodity NVIDIA GPUs using microbenchmarks, and also implement a complex image processing application that constructs a photo collage from a subset of 10 million images stored in a 40GB file. The GPU implementation maps the entire file into GPU memory and accesses it via active pointers. The use of active pointers adds only up to 1% to the application's runtime, while enabling speedups of up to 3.9× over a combined CPU+GPU implementation and 2.6× over a 12-core CPU-only implementation which uses AVX vector instructions. Sagi Shahar, Shai Bergman, Mark Silberstein |
ISCA | 2 |