EDBT 2026 Demo / reviewers in the wild / expert
Mark Silberstein
dblp:94/2996
· DBLP profile ↗
71ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0001-9659-068XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 45 · 10 first-author · 16 since 2021Computer networks · 11 · 9 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 4 since 2021Security and privacy · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In Link We Trust: BFT at the Speed of CFT using Switches
Lior Zeno, Naama Ben-David, Mark Silberstein |
NSDI | 3 |
| 2025 | AMuLeT: Automated Design-Time Testing of Secure Speculation CountermeasuresabstractIn recent years, several hardware-based countermeasures proposed to mitigate Spectre attacks have been shown to be insecure. To enable the development of effective secure speculation countermeasures, we need easy-to-use tools that can automatically test their security guarantees early-on in the design phase to facilitate rapid prototyping. Leo Tenenbaum, David Adler, Assaf Klein, Arpit Gogia, Alaa R. Alameldeen, Marco Guarnieri, Mark Silberstein, Oleksii Oleksenko, Gururaj Saileshwar |
ASPLOS (2) | 8 |
| 2025 | FlowPulse: Catching Network Failures in ML ClustersabstractNetwork hardware faults are inevitable in massive scale-out ML training clusters. Networks in such systems are inherently designed for resiliency, routing around faulty components as long as a fault is detected. Unfortunately, some silent faults evade detection. Notably, the effects of silent faults are amplified in modern production networks that deploy per-packet load balancing, because packets of a single flow traverse many network paths, making such faults particularly hard to localize. Jakob Krebs, Dimitry Gavrilenko, Daniel Amir, Shir Landau Feibish, Mark Silberstein |
HotNets | 5 |
| 2025 | Accelerating Nested Virtualization with HyperTurtle
Ori Ben Zur, Jakob Krebs, Shai Bergman, Mark Silberstein |
USENIX ATC | 4 |
| 2025 | Introduction to the Special Section on EuroSys 2024
Bianca Schroeder, Mark Silberstein |
ACM Trans. Comput. Syst. | 2 |
| 2024 | Multitenant In-Network Acceleration with SwitchVM
Sajy Khashab, Alon Rashelbach, Mark Silberstein |
NSDI | 3 |
| 2024 | In-Network Address Caching for Virtual NetworksabstractPacket routing in virtual networks requires virtual-to-physical address translation. The address mappings are updated by a single party, i.e., the network administrator, but they are read by multiple devices across the network when routing tenant packets. Existing approaches face an inherent read-write performance tradeoff: they either store these mappings in dedicated gateways for fast updates at the cost of slower forwarding or replicate them at end-hosts and suffer from slow updates. Lior Zeno, Ang Chen 0001, Mark Silberstein |
SIGCOMM | 3 |
| 2024 | Space-efficient FTL for Mobile Storage via Tiny Neural NetsabstractWe present RQFTL, a demand-based FTL for mobile storage controllers that boosts the effective Logical-To-Physical (L2P) address translation cache capacity over state-of-the-art techniques. RQFTL stores a large part of the L2P cache in a compressed form, and employs a learned data structure called RQRMI that leverages tiny neural nets to quickly find the correct translation entry in the cache. RQFTL uses neural network inference for cache lookups, and rapidly retrains the neural nets to efficiently handle L2P cache updates. It is specifically optimized to achieve high coverage for scattered read accesses, making it suitable for popular read-skewed workloads such as mobile gaming. Ron Marcus, Alon Rashelbach, Ori Ben Zur, Pavel Lifshits, Mark Silberstein |
SYSTOR | 5 |
| 2023 | NeuroLPM - Scaling Longest Prefix Match Hardware with Neural NetworksabstractLongest Prefix Match engines (LPM) are broadly used in computer systems and especially in modern network devices such as Network Interface Cards (NICs), switches and routers. However, existing LPM hardware fails to scale to millions of rules required by modern systems, is often optimized for specific applications, and thus is performance-sensitive to the structure of LPM rules. Alon Rashelbach, Igor Lima de Paula, Mark Silberstein |
MICRO | 3 |
| 2023 | Hide and Seek with Spectres: Efficient discovery of speculative information leaks with random testingabstractAttacks like Spectre abuse speculative execution, one of the key performance optimizations of modern CPUs. Recently, several testing tools have emerged to automatically detect speculative leaks in commercial (black-box) CPUs. However, the testing process is still slow, which has hindered in-depth testing campaigns, and so far prevented the discovery of new classes of leakage.In this paper, we identify the root causes of the performance limitations in existing approaches, and propose techniques to overcome these limitations. With these techniques, we improve the testing speed over the state-of-the-art by up to two orders of magnitude.These improvements enable us to run a testing campaign of unprecedented depth on Intel and AMD CPUs. As a highlight, we discover two types of previously unknown speculative leaks (affecting string comparison and division) that have escaped previous manual and automatic analyses. Oleksii Oleksenko, Marco Guarnieri, Boris Köpf, Mark Silberstein |
SP | 4 |
| 2023 | Fuzzing LibraryOSes for Iago vulnerabilitiesabstractWe present a new fuzzing approach for Iago vulnerabilities in Library OSes for SGX enclaves. Based on the filesystem model, it allows efficiently combining valid and malicious values to reach deeper paths in LibraryOS to identify more potential security vulnerabilities. Leonid Dyachkov, Meni Orenbach, Mark Silberstein |
SYSTOR | 3 |
| 2023 | SwitchVM: Multi-Tenancy for In-Network ComputingabstractWe present SwitchVM, an in-switch virtual machine for reconfigurable match-action table programmable switches aimed at providing multi-tenant in-network computing. Sajy Khashab, Mark Silberstein |
SYSTOR | 2 |
| 2023 | Neural Networks for Computer SystemsabstractWe present the Range Query Recursive Model Index (RQRMI) data structure that trades memory accesses for computations in performance-critical systems that employ Range Matching. Alon Rashelbach, Ori Rottenstreich, Mark Silberstein |
SYSTOR | 3 |
| 2023 | Reducing The Virtual Memory Overhead in Nested VirtualizationabstractVirtualization has become a critical aspect of modern computing, and with the advent of virtualization-based containers, fast nested virtualization has become increasingly important. Nested virtualization is implemented by emulating virtualization capabilities to the guest host which can result in significant overhead. Another source of overheads in virtualization stems from the address translation mechanisms employed to implement virtualization, which usually causes a mix of slower address translation, frequently trapping guests, and loss of granularity in page tables. Our research focuses on using guest-managed physical memory with the use of per-VM memory tags for checking each VMs' access permissions. Ori Ben Zur, Shai Bergman, Mark Silberstein |
SYSTOR | 3 |
| 2023 | Translation Pass-Through for Near-Native Paging Performance in VMs
Shai Bergman, Mark Silberstein, Takahiro Shinagawa, Peter R. Pietzuch, Lluís Vilanova |
USENIX ATC | 2 |
| 2023 | AEX-Notify: Thwarting Precise Single-Stepping Attacks through Interrupt Awareness for Intel SGX Enclaves
Scott Constable, Jo Van Bulck, Yuan Xiao 0001, Cedric Xing, Ilya Alexandrovich, Taesoo Kim, Frank Piessens, Mona Vij, Mark Silberstein |
USENIX Security Symposium | 10 |
| 2023 | Scaling by Learning: Accelerating Open vSwitch Data Path With Neural NetworksabstractOpen vSwitch (OVS) is a widely used open-source virtual switch implementation. In this work, we seek to scale up OVS to support hundreds of thousands of OpenFlow rules by accelerating the core component of its data-path - the packet classification mechanism. To do so we use NuevoMatch, a recent algorithm that uses neural network inference to match packets, and promises significant scalability and performance benefits. We overcome the primary algorithmic challenge of the slow training rate in the vanilla NuevoMatch, speeding it up by over three orders of magnitude. This improvement enables two design options to integrate NuevoMatch with OVS: (1) as an extra caching layer in front of OVS’s megaflow cache, and (2) using it to completely replace OVS’s data-path while performing classification directly on OpenFlow rules, and obviating control-path upcalls. Comprehensive evaluation on real-world packet traces and ClassBench rules demonstrates geometric mean speedups of$1.9\times $and$12.3\times $for the first and second designs, respectively, for 500K rules, with the latter also supporting up to 60K OpenFlow rule updates/second, by far exceeding the original OVS. Alon Rashelbach, Ori Rottenstreich, Mark Silberstein |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | ZNSwap: un-Block your SwapabstractWe introduce ZNSwap , a novel swap subsystem optimized for the recent Zoned Namespace (ZNS) SSDs. ZNSwap leverages ZNS’s explicit control over data management on the drive and introduces a space-efficient host-side Garbage Collector (GC) for swap storage co-designed with the OS swap logic. ZNSwap enables cross-layer optimizations, such as direct access to the in-kernel swap usage statistics by the GC to enable fine-grain swap storage management, and correct accounting of the GC bandwidth usage in the OS resource isolation mechanisms to improve performance isolation in multi-tenant environments. We evaluate ZNSwap using standard Linux swap benchmarks and two production key-value stores. ZNSwap shows significant performance benefits over the Linux swap on traditional SSDs, such as stable throughput for different memory access patterns, and 10× lower 99th percentile latency and 5× higher throughput for memcached key-value store under realistic usage scenarios. Shai Bergman, Niklas Cassel, Matias Bjørling, Mark Silberstein |
ACM Trans. Storage | 4 |
| 2022 | FlexDriver: a network driver for your acceleratorabstractWe propose a new system design for connecting hardware and FPGA accelerators to the network, allowing them to directly control commodity ASIC NICs without using the CPU. This solves the key challenge of leveraging existing NIC hardware offloads such as RDMA and virtualization for hardware disaggregation and accelerator networking. Our approach supports a diverse set of use cases, from direct network access for disaggregated accelerators to inline acceleration of the network stack and disaggregation of main system memory, all without implementing complex networking logic. To demonstrate this approach, we build FlexDriver (FLD) , a hardware module that implements a NIC data-plane driver. Our main technical contribution is compressing NIC control structures by \(5\times\) , allowing FLD to achieve high scalability with low die area and no memory bandwidth interference. We build two prototypes – FLD core on NVIDIA Innova-2 FPGA SmartNICs and FLD with a load/store interface on IBM OpenCAPI FPGA deployment with ConnectX-6 Dx. We demonstrate four different use cases: a disaggregated LTE cipher, an IP-reassembly inline accelerator, an IoT cryptographic-token authentication offload, and a fine-grained memory disaggregation datapath over RDMA. These leverage the ASIC NIC for RDMA processing, VXLAN tunneling, and traffic shaping, without CPU involvement. Haggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom, Noam Cohen, Amit Hermony, Dotan Levi, Liran Liss, Mark Silberstein |
ASPLOS | 9 |
| 2022 | Revizor: testing black-box CPUs against speculation contractsabstractSpeculative vulnerabilities such as Spectre and Meltdown expose speculative execution state that can be exploited to leak information across security domains via side-channels. Such vulnerabilities often stay undetected for a long time as we lack the tools for systematic testing of CPUs to find them. Oleksii Oleksenko, Christof Fetzer, Boris Köpf, Mark Silberstein |
ASPLOS | 4 |
| 2022 | Slashing the disaggregation tax in heterogeneous data centers with FractOSabstractDisaggregated heterogeneous data centers promise higher efficiency, lower total costs of ownership, and more flexibility for data-center operators. However, current software stacks can levy a high tax on application performance. Applications and OSes are designed for systems where local PCIe-connected devices are centrally managed by CPUs, but this centralization introduces unnecessary messages through the shared data-center network in a disaggregated system. Lluís Vilanova, Lina Maudlej, Shai Bergman, Till Miemietz, Matthias Hille, Nils Asmussen, Michael Roitzsch, Hermann Härtig, Mark Silberstein |
EuroSys | 9 |
| 2022 | Reconsidering OS memory optimizations in the presence of disaggregated memoryabstractTiered memory systems introduce an additional memory level with higher-than-local-DRAM access latency and require sophisticated memory management mechanisms to achieve cost-efficiency and high performance. Recent works focus on byte-addressable tiered memory architectures which offer better performance than pure swap-based systems. We observe that adding disaggregation to a byte-addressable tiered memory architecture requires important design changes that deviate from the common techniques that target lower-latency non-volatile memory systems. Our comprehensive analysis of real workloads shows that the high access latency to disaggregated memory undermines the utility of well-established memory management optimizations Based on these insights, we develop HotBox – a disaggregated memory management subsystem for Linux that strives to maximize the local memory hit rate with low memory management overhead. HotBox introduces only minor changes to the Linux kernel while outperforming state-of-the-art systems on memory-intensive benchmarks by up to 2.25×. Shai Bergman, Priyank Faldu, Boris Grot, Lluís Vilanova, Mark Silberstein |
ISMM | 5 |
| 2022 | An edge-queued datagram service for all datacenter traffic
Vladimir Andrei Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, Costin Raiciu |
NSDI | 6 |
| 2022 | Scaling Open vSwitch with a Computational Cache
Alon Rashelbach, Ori Rottenstreich, Mark Silberstein |
NSDI | 3 |
| 2022 | SwiSh: Distributed Shared State Abstractions for Programmable Switches
Lior Zeno, Dan R. K. Ports, Jacob Nelson 0001, Daehyeok Kim, Shir Landau Feibish, Idit Keidar, Arik Rinberg, Alon Rashelbach, Igor Lima de Paula, Mark Silberstein |
NSDI | 10 |
| 2022 | ZNSwap: un-Block your Swap
Shai Bergman, Niklas Cassel, Matias Bjørling, Mark Silberstein |
USENIX ATC | 4 |
| 2022 | A Computational Approach to Packet ClassificationabstractMulti-field packet classification is a crucial component in modern software-defined data center networks. To achieve high throughput and low latency, state-of-the-art algorithms strive to fit the rule lookup data structures into on-die caches; however, they do not scale well with the number of rules. We present a novel approach,NuevoMatch, which improves the memory scaling of existing methods. A new data structure,Range Query Recursive Model Index(RQ-RMI), is the key component that enables NuevoMatch to replace most of the accesses to main memory with model inference computations. We describe an efficient training algorithm that guarantees the correctness of the RQ-RMI-based classification. The use of RQ-RMI allows the rules to be compressed into neural networks that fit into the hardware cache. Further, it takes advantage of the growing support for fast neural network processing in modern CPUs, such as wide vector instructions, achieving a latency of tens of nanoseconds per lookup. Our evaluation using 500K multi-field rules from the standard ClassBench benchmark shows a geometric mean compression factor of$4.9\times $,$8\times $, and$82\times $, and average performance improvement of$2.4\times $,$2.6\times $, and$1.6\times $in throughput compared to CutSplit, NeuroCuts, and TupleMerge, all state-of-the-art algorithms. Alon Rashelbach, Ori Rottenstreich, Mark Silberstein |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | Faster Neural Network Training with Approximate Tensor OperationsabstractWe propose a novel technique for faster deep neural network training which systematically applies sample-based approximation to the constituent tensor operations, i.e., matrix multiplications and convolutions. We introduce new sampling techniques, study their theoretical properties, and prove that they provide the same convergence guarantees when applied to SGD training. We apply approximate tensor operations to single and multi-node training of MLP and CNN networks on MNIST, CIFAR-10 and ImageNet datasets. We demonstrate up to 66% reduction in the amount of computations and communication, and up to 1.37x faster training time while maintaining negligible or no impact on the final test accuracy. Menachem Adelman, Kfir Y. Levy, Ido Hakimi, Mark Silberstein |
NeurIPS | 4 |
| 2021 | Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism
Saar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein, Assaf Schuster |
USENIX ATC | 4 |
| 2020 | Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersabstractThis paper explores new opportunities afforded by the growing deployment of compute and I/O accelerators to improve the performance and efficiency of hardware-accelerated computing services in data centers. Maroun Tork, Lina Maudlej, Mark Silberstein |
ASPLOS | 3 |
| 2020 | Autarky: closing controlled channels with self-paging enclavesabstractAs the first widely-deployed secure enclave hardware, Intel SGX shows promise as a practical basis for confidential cloud computing. However, side channels remain SGX's greatest security weakness. Inparticular, the "controlled-channel attack" on enclave page faults exploits a longstanding architectural side channel and still lacks effective mitigation. Meni Orenbach, Andrew Baumann, Mark Silberstein |
EuroSys | 3 |
| 2020 | SwiShmem: Distributed Shared State Abstractions for Programmable SwitchesabstractProgrammable switches provide an appealing platform for running network functions (NFs), such as NATs, firewalls, and DDoS detectors, entirely in data plane, at staggering multi-Tbps processing rates. However, to be used in real deployments with a complex multi-switch topology, one NF instance must be deployed on each switch, which together act as a single logical NF. This requirement poses significant challenges in particular for stateful NFs, due to the need to manage distributed shared NF state among the switches. While considered a solved problem in classical distributed systems, data-plane state sharing requires addressing several unique challenges: high data rate, limited switch memory, and packet loss. Lior Zeno, Dan R. K. Ports, Jacob Nelson 0001, Mark Silberstein |
HotNets | 4 |
| 2020 | A Computational Approach to Packet ClassificationabstractMulti-field packet classification is a crucial component in modern software-defined data center networks. To achieve high throughput and low latency, state-of-the-art algorithms strive to fit the rule lookup data structures into on-die caches; however, they do not scale well with the number of rules. Alon Rashelbach, Ori Rottenstreich, Mark Silberstein |
SIGCOMM | 3 |
| 2020 | SpecFuzz: Bringing Spectre-type vulnerabilities to the surface
Oleksii Oleksenko, Bohdan Trach, Mark Silberstein, Christof Fetzer |
USENIX Security Symposium | 3 |
| 2019 | Achieving Scalability in a k-NN Multi-GPU Network Service with CentaurabstractCentaur is a GPU-centric architecture for building a low-latency approximate k-Nearest-Neighbors network server. We implement a multi-GPU distributed data flow runtime which enables efficient and scalable network request processing on GPUs. The runtime eliminates GPU management overheads from the CPU, making the server throughput and response time largely agnostic to the CPU load, speed or the number of dedicated CPU cores. Our experiments systems show that our server achieves near-perfect scaling for 16 GPUs, beating the throughput of a highly-optimized CPU-driven server by 35% while maintaining about 2msec average request latency. Furthermore, it requires only a single CPU core to run, achieving over an order of magnitude higher throughput than the standard CPU-driven server architecture in this setting. Amir Wated, Alexander Libov, Ohad Shacham, Edward Bortnikov, Mark Silberstein |
PACT | 5 |
| 2019 | Design Patterns for Code Reuse in HLS Packet Processing PipelinesabstractHigh-level synthesis (HLS) allows developers to be more productive in designing FPGA circuits thanks to familiar programming languages and high-level abstractions. In order to create high-performance circuits, HLS tools, such as Xilinx Vivado HLS, require following specific design patterns and techniques. Unfortunately, when applied to network packet processing tasks, these techniques limit code reuse and modularity, requiring developers to use deprecated programming conventions. We propose a methodology for developing high-speed networking applications using Vivado HLS for C++, focusing on reusability, code simplicity, and overall performance. Following this methodology, we implement a class library (ntl) with several building blocks that can be used in a wide spectrum of networking applications. We evaluate the methodology by implementing two applications: a UDP stateless firewall and a key-value store cache designed for FPGA-based SmartNICs, both processing packets at 40Gbps line-rate. Haggai Eran, Lior Zeno, Zsolt István, Mark Silberstein |
FCCM | 4 |
| 2019 | GAIA: An OS Page Cache for Heterogeneous Systems
Tanya Brokhman, Pavel Lifshits, Mark Silberstein |
USENIX ATC | 3 |
| 2019 | NICA: An Infrastructure for Inline Acceleration of Network Applications
Haggai Eran, Lior Zeno, Maroun Tork, Gabi Malka, Mark Silberstein |
USENIX ATC | 5 |
| 2019 | CoSMIX: A Compiler-based System for Secure Memory Instrumentation and Execution in Enclaves
Meni Orenbach, Yan Michalevsky, Christof Fetzer, Mark Silberstein |
USENIX ATC | 4 |
| 2018 | SysTEX'18: 2018 Workshop on System Software for Trusted ExecutionabstractThe rise of new processor hardware extensions that permit fine-grained and flexible trusted execution, such as Intel's SGX, ARM's TrustZone, or AMD's SEV, introduces numerous novel challenges and opportunities for developers of secure applications. There is a burning need for cross-cutting systems support of such Trusted Execution Environments (TEEs) that spans all the layers of the software stack, from the OS through runtime to compilers and programming models. The 3rd Workshop on System Software for Trusted Execution (SysTEX) will focus on systems research challenges related to TEEs, and explore new ideas and strategies for the implementation of trustworthy systems with TEEs. The workshop is also open to papers exploring attacks on current TEEs and strategies for mitigating such attacks. Baris Kasikci, Mark Silberstein |
CCS | 2 |
| 2018 | Varys: Protecting SGX Enclaves from Practical Side-Channel Attacks
Oleksii Oleksenko, Bohdan Trach, Robert Krahn, Mark Silberstein, Christof Fetzer |
USENIX ATC | 4 |
| 2018 | Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order Execution
Jo Van Bulck, Marina Minkin, Ofir Weisse, Daniel Genkin, Baris Kasikci, Frank Piessens, Mark Silberstein, Thomas F. Wenisch, Yuval Yarom, Raoul Strackx |
USENIX Security Symposium | 7 |
| 2018 | Power to peep-all: Inference Attacks by Malicious Batteries on Mobile DevicesabstractAbstract Mobile devices are equipped with increasingly smart batteries designed to provide responsiveness and extended lifetime. However, such smart batteries may present a threat to users’ privacy. We demonstrate that the phone’s power trace sampled from the battery at 1KHz holds enough information to recover a variety of sensitive information. We show techniques to infer characters typed on a touchscreen; to accurately recover browsing history in an open-world setup; and to reliably detect incoming calls, and the photo shots including their lighting conditions. Combined with a novel exfiltration technique that establishes a covert channel from the battery to a remote server via a web browser, these attacks turn the malicious battery into a stealthy surveillance device. We deconstruct the attack by analyzing its robustness to sampling rate and execution conditions. To find mitigations we identify the sources of the information leakage exploited by the attack. We discover that the GPU or DRAM power traces alone are sufficient to distinguish between different websites. However, the CPU and power-hungry peripherals such as a touchscreen are the primary sources of fine-grain information leakage. We consider and evaluate possible mitigation mechanisms, highlighting the challenges to defend against the attacks. In summary, our work shows the feasibility of the malicious battery and motivates further research into system and application-level defenses to fully mitigate this emerging threat. Pavel Lifshits, Roni Forte, Yedid Hoshen, Matthew Halpern, Manuel Philipose, Mohit Tiwari, Mark Silberstein |
Proc. Priv. Enhancing Technol. | 7 |
| 2018 | SPIN: Seamless Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUsabstractRecent GPUs enable Peer-to-Peer Direct Memory Access ( p 2 p ) from fast peripheral devices like NVMe SSDs to exclude the CPU from the data path between them for efficiency. Unfortunately, using p 2 p to access files is challenging because of the subtleties of low-level non-standard interfaces, which bypass the OS file I/O layers and may hurt system performance. Developers must possess intimate knowledge of low-level interfaces to manually handle the subtleties of data consistency and misaligned accesses. We present SPIN , which integrates p 2 p into the standard OS file I/O stack, dynamically activating p 2 p where appropriate, transparently to the user. It combines p 2 p with page cache accesses, re-enables read-ahead for sequential reads, all while maintaining standard POSIX FS consistency, portability across GPUs and SSDs, and compatibility with virtual block devices such as software RAID. We evaluate SPIN on NVIDIA and AMD GPUs using standard file I/O benchmarks, application traces, and end-to-end experiments. SPIN achieves significant performance speedups across a wide range of workloads, exceeding p 2 p throughput by up to an order of magnitude. It also boosts the performance of an aerial imagery rendering application by 2.6× by dynamically adapting to its input-dependent file access pattern, enables 3.3× higher throughput for a GPU-accelerated log server, and enables 29% faster execution for the highly optimized GPU-accelerated image collage with only 30 changed lines of code. Shai Bergman, Tanya Brokhman, Tzachi Cohen, Mark Silberstein |
ACM Trans. Comput. Syst. | 4 |
| 2017 | Computational Integrity with a Public Random String from Quasi-Linear PCPs
Eli Ben-Sasson, Iddo Bentov, Alessandro Chiesa, Ariel Gabizon, Daniel Genkin, Matan Hamilis, Evgenya Pergament, Michael Riabzev, Mark Silberstein, Eran Tromer, Madars Virza |
EUROCRYPT (3) | 9 |
| 2017 | Eleos: ExitLess OS Services for SGX EnclavesabstractIntel Software Guard extensions (SGX) enable secure and trusted execution of user code in an isolated enclave to protect against a powerful adversary. Unfortunately, running I/O-intensive, memory-demanding server applications in enclaves leads to significant performance degradation. Such applications put a substantial load on the in-enclave system call and secure paging mechanisms, which turn out to be the main reason for the application slowdown. In addition to the high direct cost of thousands-of-cycles long SGX management instructions, these mechanisms incur the high indirect cost of enclave exits due to associated TLB flushes and processor state pollution. Meni Orenbach, Pavel Lifshits, Marina Minkin, Mark Silberstein |
EuroSys | 4 |
| 2017 | OmniX: an accelerator-centric OS for omni-programmable systemsabstractFuture systemswill be omni-programmable: alongside CPUs, GPUs and FPGAs, theywill execute user code near-storage, near-network, near-memory, or on other Near-X accelerator Units, NXUs. This paper explores the design space ofOS support for omni-programmable systems, aiming to simplify the development of efficient applications that span multiple heterogeneous processors and near-data accelerators. OmniX is an accelerator-centric OS architecture that extends standard OS abstractions, such as task execution and I/O, into NXUs while maintaining a coherent viewof the systemamong all the processors. OmniX enables NXUs to directly invoke tasks and access I/O services among themselves, excluding the CPU from the performance-critical control plane operations. The host CPU serves as a controller - for protection, device configuration and monitoring.We discuss the hardware trends that motivate ourwork, outline OmniX design principles, and sketch the core implementation ideas while highlighting missing hardware features, in the hope of motivating hardware vendors to implement them soon. Mark Silberstein |
HotOS | 1 |
| 2017 | SPIN: Seamless Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs
Shai Bergman, Tanya Brokhman, Tzachi Cohen, Mark Silberstein |
USENIX ATC | 4 |
| 2016 | Optimizing distributed actor systems for dynamic interactive servicesabstractDistributed actor systems are widely used for developing interactive scalable cloud services, such as social networks and on-line games. By modeling an application as a dynamic set of lightweight communicating "actors", developers can easily build complex distributed applications, while the underlying runtime system deals with low-level complexities of a distributed environment. Andrew Newell, Gabriel Kliot, Ishai Menache, Aditya Gopalan, Soramichi Akiyama, Mark Silberstein |
EuroSys | 6 |
| 2016 | Fast Multiplication in Binary Fields on GPUs via Register CacheabstractFinite fields of characteristic 2 -- "binary fields" -- are used in a variety of applications in cryptography and data storage. Multiplication of two finite field elements is a fundamental operation and a well-known computational bottleneck in many of these applications, as they often require multiplication of a large number of elements. In this work we focus on accelerating multiplication in "large" binary fields of sizes greater than 232. We devise a new parallel algorithm optimized for execution on GPUs. This algorithm makes it possible to multiply large number of finite field elements, and achieves high performance via bit-slicing and fine-grained parallelization. Eli Ben-Sasson, Matan Hamilis, Mark Silberstein, Eran Tromer |
ICS | 3 |
| 2016 | ActivePointers: A Case for Software Address Translation on GPUsabstractModern discrete GPUs have been the processors of choice for accelerating compute-intensive applications, but using them in large-scale data processing is extremely challenging. Unfortunately, they do not provide important I/O abstractions long established in the CPU context, such as memory mapped files, which shield programmers from the complexity of buffer and I/O device management. However, implementing these abstractions on GPUs poses a problem: the limited GPU virtual memory system provides no address space management and page fault handling mechanisms to GPU developers, and does not allow modifications to memory mappings for running GPU programs. We implement ActivePointers, a software address translation layer and paging system that introduces native support for page faults and virtual address space management to GPU programs, and enables the implementation of fully functional memory mapped files on commodity GPUs. Files mapped into GPU memory are accessed using active pointers, which behave like regular pointers but access the GPU page cache under the hood, and trigger page faults which are handled on the GPU. We design and evaluate a number of novel mechanisms, including a translation cache in hardware registers and translation aggregation for deadlock-free page fault handling of threads in a single warp. We extensively evaluate ActivePointers on commodity NVIDIA GPUs using microbenchmarks, and also implement a complex image processing application that constructs a photo collage from a subset of 10 million images stored in a 40GB file. The GPU implementation maps the entire file into GPU memory and accesses it via active pointers. The use of active pointers adds only up to 1% to the application's runtime, while enabling speedups of up to 3.9× over a combined CPU+GPU implementation and 2.6× over a 12-core CPU-only implementation which uses AVX vector instructions. Sagi Shahar, Shai Bergman, Mark Silberstein |
ISCA | 3 |
| 2016 | Supporting data-driven I/O on GPUs using GPUfsabstractUsing discrete GPUs for processing very large datasets is challenging, in particular when an algorithm exhibit unpredictable, data-driven access patterns. In this paper we investigate the utility of GPUfs, a library that provides direct access to files from GPU programs, to implement such algorithms. We analyze the system's bottlenecks, and suggest several modifications to the GPUfs design, including new concurrent hash table for the buffer cache and a highly parallel memory allocator. We also show that by implementing the workload in a warp-centric manner we can improve the performance even further. We evaluate our changes by implementing a real image processing application which creates collages from a dataset of 10 Million images. The enhanced GPUfs design improves the application performance by 5.6× on average over the original GPUfs, and outperforms both 12-core parallel CPU which uses the AVX instruction set, and a standard CUDA-based GPU implementation by up to 2.5× and 3× respectively, while significantly enhancing system programmability and simplifying the application design and implementation. Sagi Shahar, Mark Silberstein |
SYSTOR | 2 |
| 2016 | GPUnet: Networking Abstractions for GPU ProgramsabstractDespite the popularity of GPUs in high-performance and scientific computing, and despite increasingly general-purpose hardware capabilities, the use of GPUs in network servers or distributed systems poses significant challenges. GPUnet is a native GPU networking layer that provides a socket abstraction and high-level networking APIs for GPU programs. We use GPUnet to streamline the development of high-performance, distributed applications like in-GPU-memory MapReduce and a new class of low-latency, high-throughput GPU-native network services such as a face verification server. Mark Silberstein, Sangman Kim, Seonggu Huh, Xinya Zhang, Yige Hu, Amir Wated, Emmett Witchel |
ACM Trans. Comput. Syst. | 1 |
| 2014 | GPUnet: Networking Abstractions for GPU Programs
Sangman Kim, Seonggu Huh, Xinya Zhang, Yige Hu, Amir Wated, Emmett Witchel, Mark Silberstein |
OSDI | 7 |
| 2014 | Lazy Means Smart: Reducing Repair Bandwidth Costs in Erasure-coded Distributed StorageabstractErasure coding schemes provide higher durability at lower storage cost, and thus constitute an attractive alternative to replication in distributed storage systems, in particular for storing rarely accessed "cold" data. These schemes, however, require an order of magnitude higher recovery bandwidth for maintaining a constant level of durability in the face of node failures. In this paper we propose lazy recovery, a technique to reduce recovery bandwidth demands down to the level of replicated storage. The key insight is that a careful adjustment of recovery rate substantially reduces recovery bandwidth, while keeping the impact on read performance and data durability low. We demonstrate the benefits of lazy recovery via extensive simulation using a realistic distributed storage configuration and published component failure parameters. For example, when applied to the commonly used RS(14, 10) code, lazy recovery reduces repair bandwidth by up to 76% even below replication, while increasing the amount of degraded stripes by 0.1 percentage points. Lazy recovery works well with a variety of erasure coding schemes, including the recently introduced bandwidth efficient codes, achieving up to a factor of 2 additional bandwidth savings. Mark Silberstein, Lakshmi Ganesh, Yang Wang 0009, Lorenzo Alvisi, Michael Dahlin |
SYSTOR | 1 |
| 2014 | GPUfs: Integrating a file system with GPUsabstractAs GPU hardware becomes increasingly general-purpose, it is quickly outgrowing the traditional, constrained GPU-as-coprocessor programming model. This article advocates for extending standard operating system services and abstractions to GPUs in order to facilitate program development and enable harmonious integration of GPUs in computing systems. As an example, we describe the design and implementation of GPUFs, a software layer which provides operating system support for accessing host files directly from GPU programs. GPUFs provides a POSIX-like API, exploits GPU parallelism for efficiency, and optimizes GPU file access by extending the host CPU's buffer cache into GPU memory. Our experiments, based on a set of real benchmarks adapted to use our file system, demonstrate the feasibility and benefits of the GPUFs approach. For example, a self-contained GPU program that searches for a set of strings throughout the Linux kernel source tree runs over seven times faster than on an eight-core CPU. Mark Silberstein, Bryan Ford, Idit Keidar, Emmett Witchel |
ACM Trans. Comput. Syst. | 1 |
| 2013 | GPUfs: integrating a file system with GPUsabstractPU hardware is becoming increasingly general purpose, quickly outgrowing the traditional but constrained GPU-as-coprocessor programming model. To make GPUs easier to program and easier to integrate with existing systems, we propose making the host's file system directly accessible from GPU code. GPUfs provides a POSIX-like API for GPU programs, exploits GPU parallelism for efficiency, and optimizes GPU file access by extending the buffer cache into GPU memory. Our experiments, based on a set of real benchmarks adopted to use our file system, demonstrate the feasibility and benefits of our approach. For example, we demonstrate a simple self-contained GPU program which searches for a set of strings in the entire tree of Linux kernel source files over seven times faster than an eight-core CPU run. Mark Silberstein, Bryan Ford, Idit Keidar, Emmett Witchel |
ASPLOS | 1 |
| 2013 | A system for exact and approximate genetic linkage analysis of SNP data in large pedigreesabstractMOTIVATION: The use of dense single nucleotide polymorphism (SNP) data in genetic linkage analysis of large pedigrees is impeded by significant technical, methodological and computational challenges. Here we describe Superlink-Online SNP, a new powerful online system that streamlines the linkage analysis of SNP data. It features a fully integrated flexible processing workflow comprising both well-known and novel data analysis tools, including SNP clustering, erroneous data filtering, exact and approximate LOD calculations and maximum-likelihood haplotyping. The system draws its power from thousands of CPUs, performing data analysis tasks orders of magnitude faster than a single computer. By providing an intuitive interface to sophisticated state-of-the-art analysis tools coupled with high computing capacity, Superlink-Online SNP helps geneticists unleash the potential of SNP data for detecting disease genes. RESULTS: Computations performed by Superlink-Online SNP are automatically parallelized using novel paradigms, and executed on unlimited number of private or public CPUs. One novel service is large-scale approximate Markov Chain-Monte Carlo (MCMC) analysis. The accuracy of the results is reliably estimated by running the same computation on multiple CPUs and evaluating the Gelman-Rubin Score to set aside unreliable results. Another service within the workflow is a novel parallelized exact algorithm for inferring maximum-likelihood haplotyping. The reported system enables genetic analyses that were previously infeasible. We demonstrate the system capabilities through a study of a large complex pedigree affected with metabolic syndrome. AVAILABILITY: Superlink-Online SNP is freely available for researchers at http://cbl-hap.cs.technion.ac.il/superlink-snp. The system source code can also be downloaded from the system website. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mark Silberstein, Omer Weissbrod, Lars Otten, Anna Tzemach, Andrei Anisenia, Oren Shtark, Dvir Tuberg, Eddie Galfrin, Irena Gannon, Adel Shalata, Zvi U. Borochowitz, Rina Dechter, Elizabeth Thompson, Dan Geiger |
Bioinform. | 1 |
| 2013 | A system for exact and approximate genetic linkage analysis of SNP data in large pedigreesabstractVol. 29, No. 2, 2013, pp. 197–205 doi:10.1093/bioinformatics/bts658 The publishers regret that the author affiliations for this paper should appear as follows: Mark Silberstein1,2, Omer Weissbrod1,*, Lars Otten3, Anna Tzemach1, Andrei Anisenia1,4, Oren Shtark1, Dvir Tuberg1, Eddie Galfrin1, Irena Gannon1, Adel Shalata5,6,7, Zvi U. Borochowitz5,8, Rina Dechter3, Elizabeth Thompson9 and Dan Geiger1 1Department of Computer Science, Technion-Israel Institute of Technology, Haifa, Israel, 2Department of Computer Science, University of Texas at Austin, Austin, TX, USA, 3Donald Bren School of Information and Computer Sciences, UC Irvine, CA, USA, 4Department of Computer Science, University of Ottawa, Ottawa, Canada, 5The Simon Winter Institute for Human Genetics, Bnai-Zion Medical Center, Haifa, Israel, 6Research and Development Center, The Galilee Society, Shefa-Amr, Israel, 7Holy Family Hospital, Nazareth, Israel, 8The Rappaport Faculty of Medicine and Research Institute, Technion-Israel Institute of Technology, Haifa, Israel and 9Department of Statistics, University of Washington, Seattle, WA, USA Mark Silberstein, Omer Weissbrod, Lars Otten, Anna Tzemach, Andrei Anisenia, Oren Shtark, Dvir Tuberg, Eddie Galfrin, Irena Gannon, Adel Shalata, Zvi U. Borochowitz, Rina Dechter, Elizabeth Thompson, Dan Geiger |
Bioinform. | 1 |
| 2012 | ExPERT: Pareto-Efficient Task Replication on Grids and a CloudabstractMany scientists perform extensive computations by executing large bags of similar tasks (BoTs) in mixtures of computational environments, such as grids and clouds. Although the reliability and cost may vary considerably across these environments, no tool exists to assist scientists in the selection of environments that can both fulfill deadlines and fit budgets. To address this situation, we introduce the Expert BoT scheduling framework. Our framework systematically selects from a large search space the Pareto-efficient scheduling strategies, that is, the strategies that deliver the best results for both make span and cost. Expert chooses from them the best strategy according to a general, user-specified utility function. Through simulations and experiments in real production environments, we demonstrate that Expert can substantially reduce both make span and cost in comparison to common scheduling strategies. For bioinformatics BoTs executed in a real mixed grid + cloud environment, we show how the scheduling strategy selected by Expert reduces both make span and cost by 30%-70%, in comparison to commonly-used scheduling strategies. Orna Agmon Ben-Yehuda, Assaf Schuster, Artyom Sharov, Mark Silberstein, Alexandru Iosup |
IPDPS | 4 |
| 2012 | Eternal Sunshine of the Spotless Machine: Protecting Privacy with Ephemeral Channels
Alan M. Dunn, Michael Z. Lee, Suman Jana, Sangman Kim, Mark Silberstein, Yuanzhong Xu, Vitaly Shmatikov, Emmett Witchel |
OSDI | 5 |
| 2012 | Scheduling processing of real-time data streams on heterogeneous multi-GPU systemsabstractProcessing vast numbers of data streams is a common problem in modern computer systems and is known as the "online big data problem." Adding hard real-time constraints to the processing makes the scheduling problem a very challenging task that this paper aims to address. In such an environment, each data stream is manipulated by a (different) application and each datum (data packet) needs to be processed within a known deadline from the time it was generated. This work assumes a central compute engine which consists of a set of CPUs and a set of GPUs. The system receives a configuration of multiple incoming streams and executes a scheduler on the CPU side. The scheduler decides where each data stream will be manipulated (on the CPUs or on one of the GPUs), and the order of execution, in a way that guarantees that no deadlines will be missed. Our scheduler finds such schedules even for workloads that require high utilization of the entire system (CPUs and GPUs). Uri Verner, Assaf Schuster, Mark Silberstein, Avi Mendelson |
SYSTOR | 3 |
| 2011 | Building an Online Domain-Specific Computing Service over Non-dedicated Grid and Cloud Resources: The Superlink-Online ExperienceabstractLinkage analysis is a statistical method used by geneticists in everyday practice for mapping disease-susceptibility genes in the study of complex diseases. An essential first step in the study of genetic diseases, linkage computations may require years of CPU time. The recent DNA sampling revolution enabled unprecedented sampling density, but made the analysis even more computationally demanding. In this paper we describe a high performance online service for genetic linkage analysis, called Super link-online. The system enables anyone with Internet access to submit genetic data and analyze it as easily and quickly as if using a supercomputer. The analyses are automatically parallelized and executed on tens of thousands distributed CPUs in multiple clouds and grids. The first version of the system, which employed up to 3,000 CPUs in UW Madison and Technion Condor pools, has been successfully used since 2006 by hundreds of geneticists worldwide, with over 40 citations in the genetics literature. Here we describe the second version, which substantially improves the scalability and performance of first: it uses over 45,000 non-dedicated hosts, in 10different grids and clouds, including EC2 and the Superlink@Technion community grid. Improved system performance is obtained through a virtual grid hierarchy with dynamic load balancing and multi-grid overlay via the Grid Bot system, parallel pruning of short tasks for overhead minimization, and cost-efficient use of cloud resources in reliability-critical execution periods. These enhancements enabled execution of many previously infeasible analyses, which can now be completed within a few hours. The new version of the system, in production since 2009, has completed over 6500 different runs of over 10 million tasks, with total consumption of 420 CPU years. Mark Silberstein |
CCGRID | 1 |
| 2011 | Processing data streams with hard real-time constraints on heterogeneous systemsabstractData stream processing applications such as stock exchange data analysis, VoIP streaming, and sensor data processing pose two conflicting challenges: short per-stream latency -- to satisfy the milliseconds-long, hard real-time constraints of each stream, and high throughput -- to enable efficient processing of as many streams as possible. High-throughput programmable accelerators such as modern GPUs hold high potential to speed up the computations. However, their use for hard real-time stream processing is complicated by slow communications with CPUs, variable throughput changing non-linearly with the input size, and weak consistency of their local memory with respect to CPU accesses. Furthermore, their coarse grain hardware scheduler renders them unsuitable for unbalanced multi-stream workloads. Uri Verner, Assaf Schuster, Mark Silberstein |
ICS | 3 |
| 2011 | PTask: operating system abstractions to manage GPUs as compute devicesabstractWe propose a new set of OS abstractions to support GPUs and other accelerator devices as first class computing resources. These new abstractions, collectively called the PTask API, support a dataflow programming model. Because a PTask graph consists of OS-managed objects, the kernel has sufficient visibility and control to provide system-wide guarantees like fairness and performance isolation, and can streamline data movement in ways that are impossible under current GPU programming models. Christopher J. Rossbach, Jon Currey, Mark Silberstein, Baishakhi Ray, Emmett Witchel |
SOSP | 3 |
| 2011 | An exact algorithm for energy-efficient acceleration of task trees on CPU/GPU architecturesabstractWe consider the problem of energy-efficient acceleration of applications comprising multiple interdependent tasks forming a dependency tree, on a hypothetical CPU/GPU system where both a CPU and a GPU can be powered off when idle. Each task in the tree can be invoked on either a GPU or a CPU, but the performance may vary: some run faster on a GPU, while others prefer a CPU, making the choice of the lowest-energy processor input dependent. Furthermore, greedily minimizing the energy consumption for each task is suboptimal because of the additional energy required for the communication between the tasks executed on different processors. Mark Silberstein, Naoya Maruyama |
SYSTOR | 1 |
| 2009 | GridBot: execution of bags of tasks in multiple gridsabstractWe present a holistic approach for efficient execution of bags-of-tasks (BOTs) on multiple grids, clusters, and volunteer computing grids virtualized as a single computing platform. The challenge is twofold: to assemble this compound environment and to employ it for execution of a mixture of throughput- and performance-oriented BOTs, with a dozen to millions of tasks each. Our generic mechanism allows per BOT specification of dynamic arbitrary scheduling and replication policies as a function of the system state, BOT execution state, and BOT priority. Mark Silberstein, Artyom Sharov, Dan Geiger, Assaf Schuster |
SC | 1 |
| 2008 | Quasi-opportunistic Supercomputing in Grid Environments
Valentin Kravtsov, David Carmeli, Werner Dubitzky, Ariel Orda, Assaf Schuster, Mark Silberstein, Benny Yoshpa |
ICA3PP | 6 |
| 2008 | Efficient computation of sum-products on GPUs through software-managed cacheabstractWe present a technique for designing memory-bound algorithms with high data reuse on Graphics Processing Units (GPUs) equipped with close-to-ALU software-managed memory. The approach is based on the efficient use of this memory through the implementation of a software-managed cache. We also present an analytical model for performance analysis of such algorithms. Mark Silberstein, Assaf Schuster, Dan Geiger, Anjul Patney, John D. Owens |
ICS | 1 |
| 2006 | Scheduling Mixed Workloads in Multi-grids: The Grid Execution HierarchyabstractConsider a workload in which massively parallel tasks that require large resource pools are interleaved with short tasks that require fast response but consume fewer resources. We aim at achieving high throughput and short response time when scheduling such a workload over a set of uncoordinated grids of varying sizes and performance characteristics. We propose the concept of a grid execution hierarchy, where available grids are sorted according to their size, and the execution overheads increase with the size of the grids. We devise a scheduling algorithm for this execution hierarchy of grids by adapting the multilevel feedback queue approach to a multi-grid environment. The algorithm finds a grid of the size, availability, and overhead that best matches a task's resource requirements and expected turnaround time. Our approach is inspired by the shortest processing time first policy (SPTF), in the sense that the task's processing demands are constantly reevaluated during its run, so that a task is migrated to a more suitable level of the execution hierarchy when appropriate. We evaluate our approach in the context of the superlink-online system for processing genetic linkage analysis tasks - a production system consisting of several grids and utilizing tens of thousands of CPU hours a month. With our approach the system provides nearly interactive response time for shorter tasks, while simultaneously serving throughput-oriented massively parallel tasks in an efficient manner Mark Silberstein, Dan Geiger, Assaf Schuster, Miron Livny |
HPDC | 1 |
| 2006 | Materializing Highly Available GridsabstractGrids are becoming a mission-critical component in research and industry. The services they provide are thus required to be highly available, contributing to the vision of the grid as a dependable virtual computer of infinite power. However, building highly available services in grid is particularly difficult due to the unique characteristics of the grid environment. We believe that high availability functionality should itself be provided as a service, which can be used by transparently decorating, but not changing, the original services, thus making them highly available. In this work we highlight the major challenges and describe our initial experience in building such a generic high availability service in the context of the Condor system Mark Silberstein, Gabriel Kliot, Artyom Sharov, Assaf Schuster, Miron Livny |
HPDC | 1 |