VLDB 2026 Research / reviewers in the wild / expert
Liran Liss
dblp:17/179
· DBLP profile ↗
12ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0003-2661-7906ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Device-Assisted Live Migration of RDMA DevicesabstractRecently, we have seen growing pressure to move highperformance workloads, such as HPC and AI, to cloud environments that offer more affordable and manageable infrastructure. These workloads require direct access to RDMA devices for high-performance communication. Device passthrough, however, violates the decoupling between the guest OS and the underlying hardware, making Live Migration (LM) extremely challenging [29, 38, 40, 42, 48]. Artem Y. Polyakov, Gal Shalom, Asaf Schwartz, Aviad Yehezkel, Omri Ben David, Omri Kahalon, Ariel Shahar, Liran Liss |
SOSP | 8 |
| 2022 | FlexDriver: a network driver for your acceleratorabstractWe propose a new system design for connecting hardware and FPGA accelerators to the network, allowing them to directly control commodity ASIC NICs without using the CPU. This solves the key challenge of leveraging existing NIC hardware offloads such as RDMA and virtualization for hardware disaggregation and accelerator networking. Our approach supports a diverse set of use cases, from direct network access for disaggregated accelerators to inline acceleration of the network stack and disaggregation of main system memory, all without implementing complex networking logic. To demonstrate this approach, we build FlexDriver (FLD) , a hardware module that implements a NIC data-plane driver. Our main technical contribution is compressing NIC control structures by \(5\times\) , allowing FLD to achieve high scalability with low die area and no memory bandwidth interference. We build two prototypes – FLD core on NVIDIA Innova-2 FPGA SmartNICs and FLD with a load/store interface on IBM OpenCAPI FPGA deployment with ConnectX-6 Dx. We demonstrate four different use cases: a disaggregated LTE cipher, an IP-reassembly inline accelerator, an IoT cryptographic-token authentication offload, and a fine-grained memory disaggregation datapath over RDMA. These leverage the ASIC NIC for RDMA processing, VXLAN tunneling, and traffic shaping, without CPU involvement. Haggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom, Noam Cohen, Amit Hermony, Dotan Levi, Liran Liss, Mark Silberstein |
ASPLOS | 8 |
| 2022 | The benefits of general-purpose on-NIC memoryabstractWe propose to use the small, newly available on-NIC memory ("nicmem") to keep pace with the rapidly increasing performance of NICs. We motivate our proposal by accelerating two types of workload classes: NFV and key-value stores. As NFV workloads frequently operate on headers---rather than data---of incoming packets, we introduce a new packet-processing architecture that splits between the two, keeping the data on nicmem when possible and thus reducing PCIe traffic, memory bandwidth, and CPU processing time. Our approach consequently shortens NFV latency by up to 23% and increases its throughput by up to 19%. Similarly, because key-value stores commonly exhibit skewed distributions, we introduce a new network stack mechanism that lets applications keep frequently accessed items on nicmem. Our design shortens memcached latency by up to 43% and increases its throughput by up to 80%. Boris Pismenny, Liran Liss, Adam Morrison 0001, Dan Tsafrir |
ASPLOS | 2 |
| 2021 | Autonomous NIC offloadsabstractCPUs routinely offload to NICs network-related processing tasks like packet segmentation and checksum. NIC offloads are advantageous because they free valuable CPU cycles. But their applicability is typically limited to layer≤4 protocols (TCP and lower), and they are inapplicable to layer-5 protocols (L5Ps) that are built on top of TCP. This limitation is caused by a misfeature we call ”offload dependence,” which dictates that L5P offloading additionally requires offloading the underlying layer≤4 protocols and related functionality: TCP, IP, firewall, etc. The dependence of L5P offloading hinders innovation, because it implies hard-wiring the complicated, ever-changing implementation of the lower-level protocols. Boris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss, Adam Morrison 0001, Dan Tsafrir |
ASPLOS | 4 |
| 2020 | IOctopus: Outsmarting Nonuniform DMAabstractIn a multi-CPU server, memory modules are local to the CPU to which they are connected, forming a nonuniform memory access (NUMA) architecture. Because non-local accesses are slower than local accesses, the NUMA architecture might degrade application performance. Similar slowdowns occur when an I/O device issues nonuniform DMA (NUDMA) operations, as the device is connected to memory via a single CPU. NUDMA effects therefore degrade application performance similarly to NUMA effects. Igor Smolyar, Alex Markuze, Boris Pismenny, Haggai Eran, Gerd Zellweger, Austin Bolen, Liran Liss, Adam Morrison 0001, Dan Tsafrir |
ASPLOS | 7 |
| 2019 | Storm: a fast transactional dataplane for remote data structuresabstractRDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance. Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera |
SYSTOR | 8 |
| 2017 | Page Fault Support for Network ControllersabstractDirect network I/O allows network controllers (NICs) to expose multiple instances of themselves, to be used by untrusted software without a trusted intermediary. Direct I/O thus frees researchers from legacy software, fueling studies that innovate in multitenant setups. Such studies, however, overwhelmingly ignore one serious problem: direct memory accesses (DMAs) of NICs disallow page faults, forcing systems to either pin entire address spaces to physical memory and thereby hinder memory utilization, or resort to APIs that pin/unpin memory buffers before/after they are DMAed, which complicates the programming model and hampers performance. Ilya Lesokhin, Haggai Eran, Shachar Raindel, Guy Shapiro, Sagi Grimberg, Liran Liss, Muli Ben-Yehuda, Nadav Amit, Dan Tsafrir |
ASPLOS | 6 |
| 2006 | Veracity radius: capturing the locality of distributed computationsabstractThis paper focuses on local computations of distributed aggregation problems on fixed graphs. We define a new metric on problem instances, Veracity Radius (VR), which captures the inherent possibility to compute them locally. We prove that VR yields a tight lower bound on output-stabilization time, i.e., the time until all nodes fix their outputs, as well as a lower bound on quiescence time. We present an efficient aggregation algorithm, I-LEAG, which reaches both output stabilization and quiescence within a time that is proportional to the VR of the problem instance, and is also efficient in terms of per-node communication and memory. We empirically show that the VR metric also effectively captures the performance of previously suggested efficient aggregation protocols, and that I-LEAG significantly outperforms these protocols in several respects. Yitzhak Birk, Idit Keidar, Liran Liss, Assaf Schuster, Ran Wolff 0001 |
PODC | 3 |
| 2006 | Efficient Dynamic Aggregation
Yitzhak Birk, Idit Keidar, Liran Liss, Assaf Schuster |
DISC | 3 |
| 2005 | GWiQ-P: an efficient decentralized grid-wide quota enforcement protocolabstractMega grids span several continents and may consist of millions of nodes and billions of tasks executing at any point in time. This setup calls for scalable and highly available resource utilization control that adapts itself to dynamic changes in the grid environment as they occur. In this paper, we address the problem of enforcing upper bounds on the consumption of grid resources. We propose a grid-wide quota enforcement system, called GWiQ-P. GWiQ-P is light-weight, and in practice is infinitely scalable, satisfying concurrently any number of resource demands, all within the limits of a global quota assigned to each user. GWiQ-P adapts to dynamic changes in the grid as they occur, improving future performance by means of improved locality. This improved performance does not impair the system's ability to respond to current requests, tolerate failures, or maintain the allotted quota levels. Kfir Karmon, Liran Liss, Assaf Schuster |
HPDC | 2 |
| 2005 | In-Kernel Integration of Operating System and Infiniband Functions for High Performance Computing Clusters: A DSM ExampleabstractThe infiniband (IB) system area network (SAN) enables applications to access hardware directly from user level, reducing the overhead of user-kernel crossings during data transfer. However, distributed applications that exhibit close coupling between network and OS services may benefit from accessing IB from the kernel through IB's native verbs interface, which permits tight integration of these services. We assess this approach using a sequential-consistency distributed shared memory (DSM) system as an example. We first develop primitives that abstract the low-level communication and kernel details, and efficiently serve the application's communication, memory, and scheduling needs. Next, we combine the primitives to form a kernel DSM protocol. The approach is evaluated using our full-fledged Linux kernel DSM implementation over infiniband. We show that overheads are reduced substantially, and overall application performance is improved in terms of both absolute execution time and scalability relative to an entirely user level implementation. Liran Liss, Yitzhak Birk, Assaf Schuster |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | A Local Algorithm for Ad Hoc Majority Voting via Charge Fusion
Yitzhak Birk, Liran Liss, Assaf Schuster, Ran Wolff 0001 |
DISC | 2 |