Haggai Eran

dblp:25/3649 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0002-2159-9046ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Semantics of Remote Direct Memory Access: Operational and Declarative Models of RDMA on TSO Architectures
abstract
Remote direct memory access (RDMA) is a modern technology enabling networked machines to exchange information without involving the operating system of either side, and thus significantly speeding up data transfer in computer clusters. While RDMA is extensively used in practice and studied in various research papers, a formal underlying model specifying the allowed behaviours of concurrent RDMA programs running in modern multicore architectures is still missing. This paper aims to close this gap and provide semantic foundations of RDMA on x86-TSO machines. We propose three equivalent formal models, two operational models in different levels of abstraction and one declarative model, and prove that the three characterisations are equivalent. To gain confidence in the proposed semantics, the more concrete operational model has been reviewed by NVIDIA experts, a major vendor of RDMA systems, and we have empirically validated the declarative formalisation on various subtle litmus tests by extensive testing. We believe that this work is a necessary initial step for formally addressing RDMA-based systems by proposing language-level models, verifying their mapping to hardware, and developing reasoning techniques for concurrent RDMA programs.
Guillaume Ambal, Brijesh Dongol, Haggai Eran, Vasileios Klimis, Ori Lahav 0001, Azalea Raad
Proc. ACM Program. Lang.3
2022 FlexDriver: a network driver for your accelerator
abstract
We propose a new system design for connecting hardware and FPGA accelerators to the network, allowing them to directly control commodity ASIC NICs without using the CPU. This solves the key challenge of leveraging existing NIC hardware offloads such as RDMA and virtualization for hardware disaggregation and accelerator networking. Our approach supports a diverse set of use cases, from direct network access for disaggregated accelerators to inline acceleration of the network stack and disaggregation of main system memory, all without implementing complex networking logic. To demonstrate this approach, we build FlexDriver (FLD) , a hardware module that implements a NIC data-plane driver. Our main technical contribution is compressing NIC control structures by \(5\times\) , allowing FLD to achieve high scalability with low die area and no memory bandwidth interference. We build two prototypes – FLD core on NVIDIA Innova-2 FPGA SmartNICs and FLD with a load/store interface on IBM OpenCAPI FPGA deployment with ConnectX-6 Dx. We demonstrate four different use cases: a disaggregated LTE cipher, an IP-reassembly inline accelerator, an IoT cryptographic-token authentication offload, and a fine-grained memory disaggregation datapath over RDMA. These leverage the ASIC NIC for RDMA processing, VXLAN tunneling, and traffic shaping, without CPU involvement.
Haggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom, Noam Cohen, Amit Hermony, Dotan Levi, Liran Liss, Mark Silberstein
ASPLOS1
2022 An edge-queued datagram service for all datacenter traffic
Vladimir Andrei Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, Costin Raiciu
NSDI2
2021 Autonomous NIC offloads
abstract
CPUs routinely offload to NICs network-related processing tasks like packet segmentation and checksum. NIC offloads are advantageous because they free valuable CPU cycles. But their applicability is typically limited to layer≤4 protocols (TCP and lower), and they are inapplicable to layer-5 protocols (L5Ps) that are built on top of TCP. This limitation is caused by a misfeature we call ”offload dependence,” which dictates that L5P offloading additionally requires offloading the underlying layer≤4 protocols and related functionality: TCP, IP, firewall, etc. The dependence of L5P offloading hinders innovation, because it implies hard-wiring the complicated, ever-changing implementation of the lower-level protocols.
Boris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss, Adam Morrison 0001, Dan Tsafrir
ASPLOS2
2020 IOctopus: Outsmarting Nonuniform DMA
abstract
In a multi-CPU server, memory modules are local to the CPU to which they are connected, forming a nonuniform memory access (NUMA) architecture. Because non-local accesses are slower than local accesses, the NUMA architecture might degrade application performance. Similar slowdowns occur when an I/O device issues nonuniform DMA (NUDMA) operations, as the device is connected to memory via a single CPU. NUDMA effects therefore degrade application performance similarly to NUMA effects.
Igor Smolyar, Alex Markuze, Boris Pismenny, Haggai Eran, Gerd Zellweger, Austin Bolen, Liran Liss, Adam Morrison 0001, Dan Tsafrir
ASPLOS4
2019 Design Patterns for Code Reuse in HLS Packet Processing Pipelines
abstract
High-level synthesis (HLS) allows developers to be more productive in designing FPGA circuits thanks to familiar programming languages and high-level abstractions. In order to create high-performance circuits, HLS tools, such as Xilinx Vivado HLS, require following specific design patterns and techniques. Unfortunately, when applied to network packet processing tasks, these techniques limit code reuse and modularity, requiring developers to use deprecated programming conventions. We propose a methodology for developing high-speed networking applications using Vivado HLS for C++, focusing on reusability, code simplicity, and overall performance. Following this methodology, we implement a class library (ntl) with several building blocks that can be used in a wide spectrum of networking applications. We evaluate the methodology by implementing two applications: a UDP stateless firewall and a key-value store cache designed for FPGA-based SmartNICs, both processing packets at 40Gbps line-rate.
Haggai Eran, Lior Zeno, Zsolt István, Mark Silberstein
FCCM1
2019 Storm: a fast transactional dataplane for remote data structures
abstract
RDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance.
Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera
SYSTOR6
2019 NICA: An Infrastructure for Inline Acceleration of Network Applications
Haggai Eran, Lior Zeno, Maroun Tork, Gabi Malka, Mark Silberstein
USENIX ATC1
2017 Page Fault Support for Network Controllers
abstract
Direct network I/O allows network controllers (NICs) to expose multiple instances of themselves, to be used by untrusted software without a trusted intermediary. Direct I/O thus frees researchers from legacy software, fueling studies that innovate in multitenant setups. Such studies, however, overwhelmingly ignore one serious problem: direct memory accesses (DMAs) of NICs disallow page faults, forcing systems to either pin entire address spaces to physical memory and thereby hinder memory utilization, or resort to APIs that pin/unpin memory buffers before/after they are DMAed, which complicates the programming model and hampers performance.
Ilya Lesokhin, Haggai Eran, Shachar Raindel, Guy Shapiro, Sagi Grimberg, Liran Liss, Muli Ben-Yehuda, Nadav Amit, Dan Tsafrir
ASPLOS2
2015 Congestion Control for Large-Scale RDMA Deployments
abstract
Modern datacenter applications demand high throughput (40Gbps) and ultra-low latency (< 10 μs per hop) from the network, with low CPU overhead. Standard TCP/IP stacks cannot meet these requirements, but Remote Direct Memory Access (RDMA) can. On IP-routed datacenter networks, RDMA is deployed using RoCEv2 protocol, which relies on Priority-based Flow Control (PFC) to enable a drop-free network. However, PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness. To alleviates these problems, we introduce DCQCN, an end-to-end congestion control scheme for RoCEv2. To optimize DCQCN performance, we build a fluid model, and provide guidelines for tuning switch buffer thresholds, and other protocol parameters. Using a 3-tier Clos network testbed, we show that DCQCN dramatically improves throughput and fairness of RoCEv2 RDMA traffic. DCQCN is implemented in Mellanox NICs, and is being deployed in Microsoft's datacenters.
Yibo Zhu 0001, Haggai Eran, Daniel Firestone, Chuanxiong Guo, Marina Lipshteyn, Yehonatan Liron, Jitendra Padhye, Shachar Raindel, Mohamad Haj Yahia, Ming Zhang 0005
SIGCOMM2
2010 Applying Constraint Programming to Identification and Assignment of Service Professionals
Sigal Asaf, Haggai Eran, Yossi Richter, Daniel P. Connors, Donna L. Gresh, Michael J. Mcinnis
CP2
2009 Pin Assignment Using Stochastic Local Search Constraint Programming
Bella Dubrov, Haggai Eran, Ari Freund 0001, Edward F. Mark, Shyam Ramji, Timothy A. Schell
CP2
2009 Transactifying Apache's cache module
abstract
Apache is a large-scale industrial multi-process and multithreaded application, which uses lock-based synchronization. We report on our experience in modifying Apache's cache module to employ transactional memory instead of locks, a process we refer to as transactification; we are not aware of any previous efforts to transactify legacy software of such a large scale. Along the way, we learned some valuable lessons about which tools one should use, which parts of the code one should transactify and which are better left untouched, as well as on the intricacy of commit handlers. We also stumbled across weaknesses of existing software transactional memory (STM) toolkits, leading us to identify desirable features they are currently lacking. Finally, we present performance results from running Apache on a 32-core machine, showing that, there are scenarios where the performance of the STM-based version is close to that of the lock-based version. These results suggest that there are applications for which the overhead of using a software-only implementation of transactional memory is insignificant.
Haggai Eran, Ohad Lutzky, Zvika Guz, Idit Keidar
SYSTOR1