Boris Pismenny

dblp:241/5842 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0009-1513-9017ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Characterizing Bluefields' Memory Bandwidth Bottlenecks
M. P. Podles, Idelfonso Tafur Monroy, Juan Jose Vegas Olmos, Boris Pismenny
PAM4
2025 Disentangling the Dual Role of NIC Receive Rings
Boris Pismenny, Adam Morrison 0001, Dan Tsafrir
OSDI1
2023 ShRing: Networking with Shared Receive Rings
Boris Pismenny, Adam Morrison 0001, Dan Tsafrir
OSDI1
2022 The benefits of general-purpose on-NIC memory
abstract
We propose to use the small, newly available on-NIC memory ("nicmem") to keep pace with the rapidly increasing performance of NICs. We motivate our proposal by accelerating two types of workload classes: NFV and key-value stores. As NFV workloads frequently operate on headers---rather than data---of incoming packets, we introduce a new packet-processing architecture that splits between the two, keeping the data on nicmem when possible and thus reducing PCIe traffic, memory bandwidth, and CPU processing time. Our approach consequently shortens NFV latency by up to 23% and increases its throughput by up to 19%. Similarly, because key-value stores commonly exhibit skewed distributions, we introduce a new network stack mechanism that lets applications keep frequently accessed items on nicmem. Our design shortens memcached latency by up to 43% and increases its throughput by up to 80%.
Boris Pismenny, Liran Liss, Adam Morrison 0001, Dan Tsafrir
ASPLOS1
2021 Autonomous NIC offloads
abstract
CPUs routinely offload to NICs network-related processing tasks like packet segmentation and checksum. NIC offloads are advantageous because they free valuable CPU cycles. But their applicability is typically limited to layer≤4 protocols (TCP and lower), and they are inapplicable to layer-5 protocols (L5Ps) that are built on top of TCP. This limitation is caused by a misfeature we call ”offload dependence,” which dictates that L5P offloading additionally requires offloading the underlying layer≤4 protocols and related functionality: TCP, IP, firewall, etc. The dependence of L5P offloading hinders innovation, because it implies hard-wiring the complicated, ever-changing implementation of the lower-level protocols.
Boris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss, Adam Morrison 0001, Dan Tsafrir
ASPLOS1
2021 Characterizing, exploiting, and detecting DMA code injection vulnerabilities in the presence of an IOMMU
abstract
Direct memory access (DMA) renders a system vulnerable to DMA attacks, in which I/O devices access memory regions not intended for their use. Hardware input-output memory management units (IOMMU) can be used to provide protection. However, an IOMMU cannot prevent all DMA attacks because it only restricts DMA at page-level granularity, leading to sub-page vulnerabilities.
Alex Markuze, Shay Vargaftik, Gil Kupfer, Boris Pismenny, Nadav Amit, Adam Morrison 0001, Dan Tsafrir
EuroSys4
2021 Rowhammering Storage Devices
abstract
Peripheral devices like SSDs are growing more complex, to the point they are effectively small computers themselves. Our position is that this trend creates a new kind of attack vector, where untrusted software could use peripherals strictly as intended to accomplish unintended goals. To exemplify, we set out to rowhammer the DRAM component of a simplified host-side FTL, issuing regular I/O requests that manage to flip bits in a way that triggers sensitive information leakage. We conclude that such attacks might soon be feasible, and we argue that systems need principled approaches for securing peripherals against them.
Tao Zhang 0045, Boris Pismenny, Donald E. Porter, Dan Tsafrir, Aviad Zuck
HotStorage2
2020 IOctopus: Outsmarting Nonuniform DMA
abstract
In a multi-CPU server, memory modules are local to the CPU to which they are connected, forming a nonuniform memory access (NUMA) architecture. Because non-local accesses are slower than local accesses, the NUMA architecture might degrade application performance. Similar slowdowns occur when an I/O device issues nonuniform DMA (NUDMA) operations, as the device is connected to memory via a single CPU. NUDMA effects therefore degrade application performance similarly to NUMA effects.
Igor Smolyar, Alex Markuze, Boris Pismenny, Haggai Eran, Gerd Zellweger, Austin Bolen, Liran Liss, Adam Morrison 0001, Dan Tsafrir
ASPLOS3
2019 Storm: a fast transactional dataplane for remote data structures
abstract
RDMA technology enables a host to access the memory of a remote host without involving the remote CPU, improving the performance of distributed in-memory storage systems. Previous studies argued that RDMA suffers from scalability issues, because the NIC's limited resources are unable to simultaneously cache the state of all the concurrent network streams. These concerns led to various software-based proposals to reduce the size of this state by trading off performance.
Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui, Yiying Zhang 0005, Haggai Eran, Boris Pismenny, Liran Liss, Michael Wei, Dan Tsafrir, Marcos K. Aguilera
SYSTOR7