EDBT 2026 Demo / reviewers in the wild / expert
Nadav Amit
dblp:52/10949
· DBLP profile ↗
26ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0002-6643-6232ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 11 · 4 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Huge Pages for Real-World ExecutablesabstractWhile huge pages can dramatically reduce address translation overhead, their use for executable code remains limited by structural barriers in binary formats, page cache management, and loader mechanisms. Existing solutions copy code into anonymous memory, abandoning file-backed semantics and breaking cross-process sharing, memory reclamation, and debugging tools. Emerging kernel support for file-backed huge pages remains insufficient without relinking and modifying loaders, impractical for deployed and closed-source software. Nadav Amit |
ISMM | 1 |
| 2025 | Batching with End-to-End Performance EstimationabstractBatching heuristics are used in multiple layers of the TCP/IP stack, aiming to improve performance by amortizing overheads. When performance is defined as average latency and throughput, optimal batching decisions can be infeasible if application-perceived end-to-end performance is unknown, which is commonly the case in general-purpose setups. We address this problem by occasionally adding a few easily maintained counters to TCP metadata exchanges and using them to estimate end-to-end performance via Little's law. We experimentally show that these estimates are accurate when application requests can be identified by the kernel (corresponding, for example, to send system calls, packets, or some fixed number of bytes). Had these estimates been used to dynamically toggle Nagle batching, they could have extended Redis's range of sustainable throughput at tolerable latencies by nearly 2x and improved latency within this range by as much as nearly 3x. When the kernel cannot identify requests on its own, we propose that applications use a simple new interface to enlighten it, thereby ensuring accuracy. Avidan Borisov, Nadav Amit, Dan Tsafrir |
HotOS | 2 |
| 2025 | Eden: Developer-Friendly Application-Integrated Far Memory
Anil Yelam, Stewart Grant, Saarth Deshpande, Nadav Amit, Radhika Niranjan Mysore, Amy Ousterhout, Marcos K. Aguilera, Alex C. Snoeren |
NSDI | 4 |
| 2025 | DeepErr: Automatic Root-Cause Analysis of System Call FailuresabstractSystem call failures present significant challenges for operating system (OS) users, as the failures are often cryptic and difficult to diagnose due to limited error codes and missing documentation. As a result, software developers struggle to utilize system calls effectively, and power users encounter difficulties configuring the OS and resolving environment problems. Existing automatic root-cause analysis tools are inadequate, primarily due to dependence on comparative analysis, which requires similar successful executions that are often unavailable. Nadav Amit, Michael Wei |
SYSTOR | 1 |
| 2024 | Every Mapping Counts in Large Amounts: Folio Accounting
David Hildenbrand, Martin Schulz 0001, Nadav Amit |
USENIX ATC | 3 |
| 2023 | Copy-on-Pin: The Missing Piece for Correct Copy-on-WriteabstractOperating systems utilize Copy-on-Write (COW) to conserve memory and improve performance. During the last two decades, a series of COW-related bugs - which compromised security, corrupted memory and degraded performance - was found. The majority of these bugs are related to page "pinning", which operating systems employ to access process memory efficiently and to perform direct I/O. Unfortunately, the true cause of these bugs is not well understood, resulting in incomplete bug fixes. We show this by: (1) surveying previously reported pinning-related COW bugs; (2) uncovering new such bugs in Linux, FreeBSD, and NetBSD; and (3) showing that they occur because the COW logic does not consider page pinnings correctly, resulting in incorrect behavior (e.g., I/O of stale data). We then address the underlying problem by deriving when/how shared pages must be copied and under which conditions pinned pages can be shared to maintain correctness. Based on this assessment, we introduce the "Copy-on-Pin (COP)" scheme, an extension of the COW mechanism that handles pinned pages correctly by ensuring pinned pages and shared pages are mutually exclusive. However, we find that a naive implementation of this scheme hampers performance and increases complexity if pages are copied only when strictly necessary. To compensate, we introduce a relaxed-COP design, which does not require precise tracking of page sharing, maintains correctness without increasing complexity, and (while potentially needlessly copying pages in some corner cases) marginally improves performance. Our relaxed-COP solution has been integrated into Linux 5.19. David Hildenbrand, Martin Schulz 0001, Nadav Amit |
ASPLOS (2) | 3 |
| 2023 | Quarantine: Mitigating Transient Execution Attacks with Physical Domain IsolationabstractSince the Spectre and Meltdown disclosure in 2018, the list of new transient execution vulnerabilities that abuse the shared nature of microarchitectural resources on CPU cores has been growing rapidly. In response, vendors keep deploying “spot” (per-variant) mitigations, which have become increasingly costly when combined against all the attacks—especially on older-generation processors. Indeed, some are so expensive that system administrators may not deploy them at all. Worse still, spot mitigations can only address known (N-day) attacks as they do not tackle the underlying problem: different security domains that run simultaneously on the same physical CPU cores and share their microarchitectural resources. Mathé Hertogh, Manuel Wiesinger, Sebastian Österlund, Marius Muench, Nadav Amit, Herbert Bos, Cristiano Giuffrida |
RAID | 5 |
| 2021 | Characterizing, exploiting, and detecting DMA code injection vulnerabilities in the presence of an IOMMUabstractDirect memory access (DMA) renders a system vulnerable to DMA attacks, in which I/O devices access memory regions not intended for their use. Hardware input-output memory management units (IOMMU) can be used to provide protection. However, an IOMMU cannot prevent all DMA attacks because it only restricts DMA at page-level granularity, leading to sub-page vulnerabilities. Alex Markuze, Shay Vargaftik, Gil Kupfer, Boris Pismenny, Nadav Amit, Adam Morrison 0001, Dan Tsafrir |
EuroSys | 5 |
| 2021 | Dealing with (some of) the fallout from meltdownabstractThe meltdown vulnerability allows users to read kernel memory by exploiting a hardware flaw in speculative execution. Processor vendors recommend "page table isolation" (PTI) as a software fix, but PTI can significantly degrade the performance of system-call-heavy programs. Leveraging the fact that 32-bit pointers cannot access 64-bit kernel memory, we propose "Shrink", a safe alternative to PTI, which is applicable to programs capable of running in 32-bit address spaces. We show that Shrink can restore the performance of some workloads, suggest additional potential alternatives, and argue that vendors must be more open about hardware flaws to allow developers to design protection schemes that are safe and performant. Nadav Amit, Michael Wei, Dan Tsafrir |
SYSTOR | 1 |
| 2020 | Don't shoot down TLB shootdowns!abstractTranslation Lookaside Buffers (TLBs) are critical for building performant virtual memory systems. Because most processors do not provide coherence for TLB mappings, TLB shootdowns provide a software mechanism that invokes inter-processor interrupts (IPLs) to synchronize TLBs. TLB shootdowns are expensive, so recent work has aimed to avoid the frequency of shootdowns through techniques such as batching. We show that aggressive batching can cause correctness issues and addressing them can obviate the benefits of batching. Instead, our work takes a different approach which focuses on both improving the performance of TLB shootdowns and carefully selecting where to avoid shootdowns. We introduce four general techniques to improve shootdown performance: (1) concurrently flush initiator and remote TLBs, (2) early acknowledgement from remote cores, (3) cacheline consolidation of kernel data structures to reduce cacheline contention, and (4) in-context flushing of userspace entries to address the overheads introduced by Spectre and Meltdown mitigations. We also identify that TLB flushing can be avoiding when handling copy-on-write (CoW) faults and some TLB shootdowns can be batched in certain system calls. Overall, we show that our approach results in significant speedups without sacrificing safety and correctness in both microbenchmarks and real-world applications. Nadav Amit, Amy Tai, Michael Wei |
EuroSys | 1 |
| 2020 | RAIDP: replication with intra-disk parityabstractDistributed storage systems often triplicate data to reduce the risk of permanent data loss, thereby tolerating at least two simultaneous disk failures at the price of 2/3 of the capacity. To reduce this price, some systems utilize erasure coding. But this optimization is usually only applied to cold data, because erasure coding might hinder performance for warm data. Eitan Rosenfeld, Aviad Zuck, Nadav Amit, Michael Factor, Dan Tsafrir |
EuroSys | 3 |
| 2019 | Using SMT to accelerate nested virtualizationabstractIaaS datacenters offer virtual machines (VMs) to their clients, who in turn sometimes deploy their own virtualized environments, thereby running a VM inside a VM. This is known as nested virtualization. Lluís Vilanova, Nadav Amit, Yoav Etsion |
ISCA | 2 |
| 2019 | Simple and precise static analysis of untrusted Linux kernel extensionsabstractExtended Berkeley Packet Filter (eBPF) is a Linux subsystem that allows safely executing untrusted user-defined extensions inside the kernel. It relies on static analysis to protect the kernel against buggy and malicious extensions. As the eBPF ecosystem evolves to support more complex and diverse extensions, the limitations of its current verifier, including high rate of false positives, poor scalability, and lack of support for loops, have become a major barrier for developers. Elazar Gershuni, Nadav Amit, Arie Gurfinkel, Nina Narodytska, Jorge A. Navas, Noam Rinetzky, Leonid Ryzhyk, Shmuel Sagiv |
PLDI | 2 |
| 2019 | JumpSwitches: Restoring the Performance of Indirect Branches In the Era of Spectre
Nadav Amit, Fred Jacobs, Michael Wei |
USENIX ATC | 1 |
| 2018 | Remote regions: a simple abstraction for remote memory
Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Stanko Novakovic, Arun Ramanathan, Pratap Subrahmanyam, Lalith Suresh 0001, Kiran Tati, Rajesh Venkatasubramanian, Michael Wei |
USENIX ATC | 2 |
| 2018 | The Design and Implementation of Hyperupcalls
Nadav Amit, Michael Wei |
USENIX ATC | 1 |
| 2017 | Page Fault Support for Network ControllersabstractDirect network I/O allows network controllers (NICs) to expose multiple instances of themselves, to be used by untrusted software without a trusted intermediary. Direct I/O thus frees researchers from legacy software, fueling studies that innovate in multitenant setups. Such studies, however, overwhelmingly ignore one serious problem: direct memory accesses (DMAs) of NICs disallow page faults, forcing systems to either pin entire address spaces to physical memory and thereby hinder memory utilization, or resort to APIs that pin/unpin memory buffers before/after they are DMAed, which complicates the programming model and hampers performance. Ilya Lesokhin, Haggai Eran, Shachar Raindel, Guy Shapiro, Sagi Grimberg, Liran Liss, Muli Ben-Yehuda, Nadav Amit, Dan Tsafrir |
ASPLOS | 8 |
| 2017 | Remote memory in the age of fast networksabstractAs the latency of the network approaches that of memory, it becomes increasingly attractive for applications to use remote memory---random-access memory at another computer that is accessed using the virtual memory subsystem. This is an old idea whose time has come, in the age of fast networks. To work effectively, remote memory must address many technical challenges. In this paper, we enumerate these challenges, discuss their feasibility, explain how some of them are addressed by recent work, and indicate other promising ways to tackle them. Some challenges remain as open problems, while others deserve more study. In this paper, we hope to provide a broad research agenda around this topic, by proposing more problems than solutions. Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Pratap Subrahmanyam, Lalith Suresh 0001, Kiran Tati, Rajesh Venkatasubramanian, Michael Wei |
SoCC | 2 |
| 2017 | Hypercallbacks: Decoupling Policy Decisions and Executionabstractresearch-article Share on Hypercallbacks: Decoupling Policy Decisions and Execution Authors: Nadav Amit VMware Research Group VMware Research GroupView Profile , Michael Wei VMware Research Group VMware Research GroupView Profile , Cheng-Chun Tu VMware VMwareView Profile Authors Info & Claims HotOS '17: Proceedings of the 16th Workshop on Hot Topics in Operating SystemsMay 2017 Pages 37–41https://doi.org/10.1145/3102980.3102987Published:07 May 2017Publication History 2citation337DownloadsMetricsTotal Citations2Total Downloads337Last 12 Months21Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Nadav Amit, Michael Wei, Cheng-Chun Tu |
HotOS | 1 |
| 2017 | Optimizing the TLB Shootdown Algorithm with Page Access Tracking
Nadav Amit |
USENIX ATC | 1 |
| 2015 | rIOMMU: Efficient IOMMU for I/O Devices that Employ Ring BuffersabstractThe IOMMU allows the OS to encapsulate I/O devices in their own virtual memory spaces, thus restricting their DMAs to specific memory pages. The OS uses the IOMMU to protect itself against buggy drivers and malicious/errant devices. But the added protection comes at a cost, degrading the throughput of I/O-intensive workloads by up to an order of magnitude. This cost has motivated system designers to trade off some safety for performance, e.g., by leaving stale information in the IOTLB for a while so as to amortize costly invalidations. We observe that high-bandwidth devices---like network and PCIe SSD controllers---interact with the OS via circular ring buffers that induce a sequential, predictable workload. We design a ring IOMMU (rIOMMU) that leverages this characteristic by replacing the virtual memory page table hierarchy with a circular, flat table. A flat table is adequately supported by exactly one IOTLB entry, making every new translation an implicit invalidation of the former and thus requiring explicit invalidations only at the end of I/O bursts. Using standard networking benchmarks, we show that rIOMMU provides up to 7.56x higher throughput relative to the baseline IOMMU, and that it is within 0.77--1.00x the throughput of a system without IOMMU protection. Moshe Malka, Nadav Amit, Muli Ben-Yehuda, Dan Tsafrir |
ASPLOS | 2 |
| 2015 | Efficient Intra-Operating System Protection Against Harmful DMAs
Moshe Malka, Nadav Amit, Dan Tsafrir |
FAST | 2 |
| 2015 | Virtual CPU validationabstractTesting the hypervisor is important for ensuring the correct operation and security of systems, but it is a hard and challenging task. We observe, however, that the challenge is similar in many respects to that of testing real CPUs. We thus propose to apply the testing environment of CPU vendors to hypervisors. We demonstrate the advantages of our proposal by adapting Intel's testing facility to the Linux KVM hypervisor. We uncover and fix 117 bugs, six of which are security vulnerabilities. We further find four flaws in Intel virtualization technology, causing a disparity between the observable behavior of code running on virtual and bare-metal servers. Nadav Amit, Dan Tsafrir, Assaf Schuster, Ahmad Ayoub, Eran Shlomo |
SOSP | 1 |
| 2014 | VSwapper: a memory swapper for virtualized environmentsabstractThe number of guest virtual machines that can be consolidated on one physical host is typically limited by the memory size, motivating memory overcommitment. Guests are given a choice to either install a "balloon" driver to coordinate the overcommitment activity, or to experience degraded performance due to uncooperative swapping. Ballooning, however, is not a complete solution, as hosts must still fall back on uncooperative swapping in various circumstances. Additionally, ballooning takes time to accommodate change, and so guests might experience degraded performance under changing conditions. Nadav Amit, Dan Tsafrir, Assaf Schuster |
ASPLOS | 1 |
| 2012 | ELI: bare-metal performance for I/O virtualizationabstractDirect device assignment enhances the performance of guest virtual machines by allowing them to communicate with I/O devices without host involvement. But even with device assignment, guests are still unable to approach bare-metal performance, because the host intercepts all interrupts, including those interrupts generated by assigned devices to signal to guests the completion of their I/O requests. The host involvement induces multiple unwarranted guest/host context switches, which significantly hamper the performance of I/O intensive workloads. To solve this problem, we present ELI (ExitLess Interrupts), a software-only approach for handling interrupts within guest virtual machines directly and securely. By removing the host from the interrupt handling path, ELI manages to improve the throughput and latency of unmodified, untrusted guests by 1.3x-1.6x, allowing them to reach 97%-100% of bare-metal performance even for the most demanding I/O-intensive workloads. Abel Gordon, Nadav Amit, Nadav Har'El, Muli Ben-Yehuda, Alex Landau, Assaf Schuster, Dan Tsafrir |
ASPLOS | 2 |
| 2011 | vIOMMU: Efficient IOMMU Emulation
Nadav Amit, Muli Ben-Yehuda, Dan Tsafrir, Assaf Schuster |
USENIX ATC | 1 |