Joel Nider

dblp:176/1026 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
4since 2021 · last 2022
0000-0001-7268-0525ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2022 Bulk JPEG decoding on in-memory processors
abstract
JPEG is a common encoding format for digital images. Applications that process large numbers of images can be accelerated by decoding multiple images concurrently. We examine the suitability of using a large array of in-memory processors (PIM) to obtain a high throughput of decoding. The main drawback of PIM processors is that they do not have the same architectural features that are commonly found on CPUs such as floating point, vector units and hardware-managed caches. Despite the lack of features, we demonstrate that it is feasible to build a JPEG decoder for PIM, and evaluate its quality and potential speedup. We show that the quality of decoded images is sufficient for real applications, and there is a significant potential for accelerating image decoding for those applications. We share our experiences in building such a decoder, and the challenges we faced while doing so.
Joel Nider, Jackson Dagger, Niloofar Gharavi, Daniel Ng, Alexandra Fedorova
SYSTOR1
2021 The last CPU
abstract
Since the end of Dennard scaling and Moore's Law have been foreseen, specialized hardware has become the focus for continued scaling of application performance. Programmable accelerators such as smart memory, smart disks, and smart NICs are now being integrated into our systems. Many accelerators can be programmed to process their data autonomously and require little or no intervention during normal operation. In this way, entire applications are offloaded, leaving the CPU with the minimal responsibilities of initialization, coordination and error handling.
Joel Nider, Alexandra Fedorova
HotOS1
2021 Jumpgate: automating integration of network connected accelerators
abstract
Network-connected accelerators (NCA), such as programmable switches, ASICs, and FPGAs can speed up operations in data analytics. But so far, integration of NCAs into data analytics systems required manual effort.
Craig Mustard, Swati Goswami, Niloofar Gharavi, Joel Nider, Ivan Beschastnikh, Alexandra Fedorova
SYSTOR4
2021 A Case Study of Processing-in-Memory in off-the-Shelf Systems
Joel Nider, Craig Mustard, Andrada Zoltan, John Ramsden, Larry Liu, Jacob Grossbard, Mohammad Dashti 0002, Romaric Jodin, Alexandre Ghiti, Jordi Chauzi, Alexandra Fedorova
USENIX ATC1
2020 Processing in Storage Class Memory
Joel Nider, Craig Mustard, Andrada Zoltan, Alexandra Fedorova
HotStorage1
2019 Unleashing the power of unikernels with unikraft
abstract
Recent research has shown that unikernels, lightweight virtual machines tailored to specific applications, have great potential in terms of performance, tiny boot times, small memory consumption, and a reduced trusted compute base. Creating and optimizing them, however, is currently a painful, time-consuming process that often needs redoing for every application. With Unikraft, we introduce a system for automatically building unikernels that drastically reduces this time without negatively impacting performance.
Simon Kuenzer, Sharan Santhanam, Yuri Volchkov, Felipe Huici, Joel Nider, Mike Rapoport, Costin Lupu
SYSTOR6
2019 Address space isolation in the linux kernel
abstract
Monolithic kernel design mandates the use of a single address space for kernel data and code. While this design is easy to understand and performs well, it does not provide much in the way of protection from exploitable bugs in the interface. By dividing up kernel objects into areas of responsibility, we can introduce additional address spaces which will prevent information leakage, even in the case of a successful attack on the kernel. We are exploring several possible implementations with the goal of increasing security while minimizing the impact on performance.
Joel Nider, Mike Rapoport, James Bottomley
SYSTOR1
2017 Zero-copy receive path in virtio
abstract
In the KVM hypervisor, incoming packets from the network must pass through several objects in the Linux kernel before being delivered to the guest VM. Currently, both the hypervisor and the guest keep their own sets of buffers on the receive path. For large packets, the overall processing time is dominated by the copying of data from hypervisor buffers to guest buffers.
Kalman Z. Meth, Mike Rapoport, Joel Nider, Razya Ladelsky
SYSTOR3
2017 Remote page faults with a CAPI based FPGA
abstract
Post-copy VM or container migration requires that the bulk of the memory is transferred after resuming on the destination node [3]. Transferring memory between nodes over a commodity TCP/IP network incurs too much latency, which slows down execution of the application on the destination node.
Joel Nider, Yiftach Binyamini, Mike Rapoport
SYSTOR1
2017 User space memory management for post-copy migration
abstract
Post-copy migration allows reduction of application down-time and reduces overall network bandwidth used for application migration [4]. Migration can be used to help optimize several aspects of operations such as power efficiency[3]. The userfault technology recently introduced to the Linux kernel allows post-copy migration of virtual machines. However, this technology is missing essential features required for post-copy migration of Linux containers.
Mike Rapoport, Joel Nider
SYSTOR2
2016 Paravirtual Remote I/O
abstract
The traditional "trap and emulate" I/O paravirtualization model conveniently allows for I/O interposition, yet it inherently incurs costly guest-host context switches. The newer "sidecore" model eliminates this overhead by dedicating host (side)cores to poll the relevant guest memory regions and react accordingly without context switching. But the dedication of sidecores on each host might be wasteful when I/O activity is low, or it might not provide enough computational power when I/O activity is high. We propose to alleviate this problem at rack scale by consolidating the dedicated sidecores spread across several hosts onto one server. The hypervisor is then effectively split into two parts: the local hypervisor that hosts the VMs, and the remote hypervisor that processes their paravirtual I/O. We call this model vRIO---paraVirtual Remote I/O. We find that by increasing the latency somewhat, it provides comparable throughput with fewer sidecores and superior throughput with the same number of sidecores as compared to the state of the art. vRIO additionally constitutes a new, cost-effective way to consolidate I/O devices (on the remote hypervisor) while supporting efficient programmable I/O interposition.
Yossi Kuperman, Eyal Moscovici, Joel Nider, Razya Ladelsky, Abel Gordon, Dan Tsafrir
ASPLOS3
2016 Workload Management for Power Efficiency in Heterogeneous Data Centers
abstract
The cloud computing paradigm has recently emerged as a convenient solution for running different workloads on highly parallel and scalable infrastructures. One major appeal of cloud computing is its capability of abstracting hardware resources and making them easy to use. Conversely, one of the major challenges for cloud providers is the energy efficiency improvement of their infrastructures. Aimed at overcoming this challenge, heterogeneous architectures have started to become part of the standard equipment used in data centers. Despite this effort, heterogeneous systems remain difficult to program and manage, while their effectiveness has been proven only in the HPC domain. Cloud workloads are different in nature and a way to exploit heterogeneity effectively is still lacking. This paper takes a first step towards an effective use of heterogeneous architectures in cloud infrastructures. It presents an in-depth analysis of cloud workloads, highlighting where energy efficiency can be obtained. The microservices paradigm is then presented as a way of intelligently partitioning applications in such a way that different components can take advantage of the heterogeneous hardware, thus providing energy efficiency. Finally, the integration of microservices and heterogeneous architectures, as well as the challenge of managing legacy applications, is presented in the context of the OPERA project.
Pietro Ruiu, Alberto Scionti, Joel Nider, Mike Rapoport
CISIS3
2016 OPERA: A Low Power Approach to the Next Generation Cloud Infrastructures
abstract
The continuous evolution of information and communication technology has led to a change in the adopted computing paradigms over time. Cloud computing is an emerging paradigm in which users, depending on their specific requirements, access to a shared pool of computing resources dynamically allocated. Cloud computing represents, with respect to Grid computing, the evolutionary step towards the implementation of a ubiquitous computing service. Such paradigm leverages on the infrastructural capabilities (compute, storage, and network) of modern data centers to provide an adequate level of computational power able to satisfy users' requests. However, trying to continuously increase such capabilities comes at the cost of an increased energy consumption. Energy efficiency is, therefore, one of the major challenges that cloud providers must address. The OPERA project aims at bringing innovative solutions to increase the energy efficiency of cloud infrastructures, by leveraging on modular, high-density, heterogeneous and low power computing systems, which are able to cover the whole computing continuum. To this end, the project will design a high-density server solution in which low power processors and FPGA devices will be used to accelerate cloud workloads. High-speed optical interconnections will be used to connect the proposed server with high-performance nodes, such as OpenPOWER-based machines. Cyber-Physical Systems (CPS) represents a natural extension of cloud infrastructures since they can collect and process data locally, more specifically where they were generated. OPERA aims at researching energy efficiency of such cloud end-nodes by designing an ultra-low power computing system with reconfigurable radio frequency capabilities. The effectiveness of the whole platform will be demonstrated with key scenarios, specifically a road traffic monitoring application, the deployment of a virtual desktop infrastructure, and the deployment of a small data center on a truck.
Alberto Scionti, Pietro Ruiu, Olivier Terzo, Joel Nider, Craig Petrie, Niccolo Baldoni
DSD4
2016 Using Storage Class Memory Efficiently for an In-memory Database
abstract
Storage class memory (SCM) is an emerging class of memory devices that are both byte addressable, and persistent. There are many different technologies that can be considered SCM, at different stages of maturity. Examples of such technologies include NVDIMM-N, PCM, SttRAM, Racetrack, FeRAM, and others.
Yonatan Gottesman, Joel Nider, Ronen I. Kat, Yaron Weinsberg, Michael Factor
SYSTOR2
2016 IO Core Manager for Virtual Environments
abstract
Para-virtualization is the leading approach in IO device virtualization. It allows the hypervisor to interpose on and inspect a virtual machine's I/O traffic at run-time. Examples of such interfaces are KVM's virtio [6] and VMWare's VMXNET [7]. Current implementations of virtual I/O in the hypervisor have been shown to have performance and scalability limitations [2, 3, 5].
Eyal Moscovici, Dan Tsafrir, Yossi Kuperman, Joel Nider, Razya Ladelsky, Abel Gordon
SYSTOR4
2016 Cross-ISA Container Migration
abstract
Containers are a convenient way of encapsulating and isolating applications. They incur less overhead than virtual machines and provide more flexibility and versatility to improve server utilization. Many new cloud applications are being written in the microservices style to take advantage of container technologies. Each component of the application can be encapsulated in a separate container, which enables the use of other features such as auto-scaling. However, legacy applications can also benefit from containers which provide more efficient development and deployment models.
Joel Nider, Mike Rapoport
SYSTOR1
2015 High performance fault-tolerance for clouds
abstract
Cloud computing and virtualized infrastructures are currently the baseline environments for the provision of services in different application domains. While the number of service consumers increasingly grows, service providers aim at exploiting infrastructures that enable non-disruptive service provisioning, thus minimizing or even eliminating downtime. Nonetheless, to achieve the latter current approaches are either application-specific or cost inefficient, requiring the use of dedicated hardware. In this paper we present the reference architecture of a fault-tolerance scheme, which not only enhances cloud environments with the aforementioned capabilities but also achieves high-performance as required by mission critical every day applications. To realize the proposed approach, a new paradigm for memory and I/O externalization and consolidation is introduced, while current implementation references are also provided.
Dimosthenis Kyriazis, Vasileios I. Anagnostopoulos, Andrea Arcangeli, Dimitrios Kalogeras, Ronen I. Kat, Cristian Klein, Panagiotis C. Kokkinos, Yossi Kuperman, Joel Nider, Petter Svärd, Luis Tomás, Emmanouel A. Varvarigos, Theodora A. Varvarigou
ISCC10