Simon Peter 0001

dblp:43/3917 · DBLP profile ↗
← Back
43ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0007-4748-8524ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 21 · 2 first-author · 7 since 2021Systems, architecture and hardware · 18 · 3 first-author · 7 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reducing the GPU Memory Bottleneck with Lossless Compression for ML
abstract
Machine learning (ML) training and inference often process data sets far exceeding GPU memory capacity, forcing them to rely on PCIe for on-demand tensor transfers, causing critical transfer bottlenecks. Lossy compression has been proposed to relieve bottlenecks but introduces workload-dependent accuracy loss, making it complex or even prohibitive to use in existing ML deployments.
Aditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon Peter 0001
EuroSys4
2026 Prediction-Informed Power Management for General-Purpose Compute Servers
abstract
This paper presents PIP (Prediction-Informed Power), a power control framework for general-purpose compute servers. PIP introduces two key innovations: (1) a machine learning-based power model that predicts the impact of hypothetical CPU throttling actions before execution, and (2) a prediction-informed control loop that selects CPU configurations to maximize performance and power utilization based on these predictions. By leveraging finegrained runtime CPU metrics, PIP can accurately estimate counterfactual power usage, allowing the control system to align power demand with the budget more quickly. Unlike traditional reactive approaches, PIP maintains effective control under frequent budget fluctuations, achieving safe oversubscription by up to 70%. Our evaluation on diverse application workloads, none of which are included in the model's training set, shows that PIP yields up to a 3.2× speedup over a state-of-the-art feedback-based system for single-application runs, and up to a 3.4× speedup for multi-application scenarios under power constraints.
Jonggyu Park, Simon Peter 0001, Thomas E. Anderson
EuroSys2
2026 PASS: A Power Adaptive Storage Server
abstract
Power management has become important in data centers. Since data center workloads are often dynamic, it is common practice to conserve energy by scaling resources up or down to match the workload. And since data centers often oversubscribe the power delivery infrastructure, operators can add power capping on top of these workload-proportional systems to adjust to available power. We find that this combination leaves performance on the table and provides only a limited power control range. Instead, we argue for power-adaptive systems that attempt to make the best use of the available power budget. To illustrate this approach, we built PASS, a power-adaptive storage system. PASS considers the interactions between different system components, including software and hardware, when making its power management decisions. For example, when under the same power constraint, PASS achieves 3–25× better throughput on filebench workloads than Intel's SPDK storage stack with Google Thunderbolt, a state of the art power capping system.
Dedong Xie, Theano Stavrinos, Jonggyu Park, Simon Peter 0001, Baris Kasikci, Thomas E. Anderson
EuroSys4
2026 Presto: A Match-Action TCP Stack for the Terabit Era
abstract
We present Presto, the first TCP stack that delivers ASIC-class performance and energy efficiency on programmable Reconfigurable Match-Action Table (RMT) pipelines, providing flexibility while retaining standard TCP semantics and POSIX socket compatibility. The key challenge in designing Presto is reconciling TCP's complex, dependent state updates with RMT's unidirectional, lock-step execution model. To overcome this challenge, Presto introduces three novel techniques: optimistic concurrency (speculative updates validated downstream), pseudo-segment injection (circular dependency resolution without stalls), and bump-in-the-wire processing (singlepass segment handling). Together, these enable TCP retransmission, reassembly, flow, and congestion control, as a pipeline of simple match-action operations.
Rajath Shashidhara, Antoine Kaufmann, Simon Peter 0001
SIGCOMM3
2025 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 0001, Ramachandran Ramjee, Ashish Panwar
ASPLOS (2)4
2024 Can Storage Devices be Power Adaptive?
abstract
Power is becoming a scarce resource for data centers, raising the need for power adaptive system design---the ability to dynamically change power consumption---to match available power. Storage makes up an increasing fraction of total data center power consumption. As such, it holds great potential to contribute to data center power adaptivity.
Dedong Xie, Theano Stavrinos, Kan Zhu, Simon Peter 0001, Baris Kasikci, Thomas E. Anderson
HotStorage4
2024 (MC)2: Lazy MemCopy at the Memory Controller
abstract
(MC)2is a lazy memory copy mechanism which can be used within memcpy-like functions to significantly reduce the CPU overhead for copies that are sparsely accessed. It can also hide copy latencies by enhancing the CPU’s ability to execute them asynchronously. (MC)2,s lazy memcpy avoids copying data at the time of invocation. Instead, (MC)2tracks prospective copies. If copied data is later accessed by a CPU or the cache, (MC)2uses the tracking information to lazily execute a copy, when necessary. Placing (MC)2at the memory controller puts it at the perfect vantage point to eliminate the largest source of memcpy overhead–CPU stalls due to cache misses in the critical path–while imposing minimal overhead itself. (MC)2consists of three main components: memory controller extensions that implement a lazy memcpy operation, a new instruction exposing the lazy memcpy, and a flexible software wrapper with semantics identical to memcpy. We implement and evaluate (MC)2in the gem5 simulator using a variety of microbenchmarks and workloads, including Google’s Protobuf, where (MC)2provides a $43 \%$ speedup and Linux huge page copy-on-write faults, where (MC)2provides $250 \times$ lower latency.
Aditya K. Kamath, Simon Peter 0001
ISCA2
2023 ScaleDB: A Scalable, Asynchronous In-Memory Database
Syed Akbar Mehdi, Deukyeon Hwang, Simon Peter 0001, Lorenzo Alvisi
OSDI3
2022 RDMA is Turing complete, we just did not know it yet!
Waleed Reda, Marco Canini, Dejan Kostic, Simon Peter 0001
NSDI4
2022 FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism
Rajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon Peter 0001
NSDI4
2022 zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IO
Tim Stamler, Deukyeon Hwang, Amanda Raybuck, Simon Peter 0001
OSDI5
2021 Rethinking File Mapping for Persistent Memory
Ian Neal, Gefei Zuo, Eric Shiple, Tanvir Ahmed Khan 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci
FAST6
2021 Rearchitecting Linux Storage Stack for µs Latency and High Throughput
Jae-Hyun Hwang, Midhul Vuppalapati, Simon Peter 0001, Rachit Agarwal 0001
OSDI3
2021 LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline Parallelism
abstract
In multi-tenant systems, the CPU overhead of distributed file systems (DFSes) is increasingly a burden to application performance. CPU and memory interference cause degraded and unstable application and storage performance, in particular for operation latency. Recent client-local DFSes for persistent memory (PM) accelerate this trend. DFS offload to SmartNICs is a promising solution to these problems, but it is challenging to fit the complex demands of a DFS onto simple SmartNIC processors located across PCIe.
Jongyul Kim 0001, Insu Jang, Waleed Reda, Jaeseong Im, Marco Canini, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Emmett Witchel
SOSP8
2021 HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM
abstract
High-capacity non-volatile memory (NVM) is a new main memory tier. Tiered DRAM+NVM servers increase total memory capacity by up to 8x, but can diminish memory bandwidth by up to 7x and inflate latency by up to 63% if not managed well. We study existing hardware and software tiered memory management systems on the recently available Intel Optane DC NVM with big data applications and find that no existing system maximizes application performance on real NVM.
Amanda Raybuck, Tim Stamler, Mattan Erez, Simon Peter 0001
SOSP5
2020 Assise: Performance and Availability via Client-local NVM in a Distributed File System
Thomas E. Anderson, Marco Canini, Jongyul Kim 0001, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Waleed Reda, Henry Schuh, Emmett Witchel
OSDI6
2020 AGAMOTTO: How Persistent is your Persistent Memory Application?
Ian Neal, Ben Reeves, Benjamin Stoler, Andrew Quinn 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci
OSDI6
2019 TAS: TCP Acceleration as an OS Service
abstract
As datacenter network speeds rise, an increasing fraction of server CPU cycles is consumed by TCP packet processing, in particular for remote procedure calls (RPCs). To free server CPUs from this burden, various existing approaches have attempted to mitigate these overheads, by bypassing the OS kernel, customizing the TCP stack for an application, or by offloading packet processing to dedicated hardware. In doing so, these approaches trade security, agility, or generality for efficiency. Neither trade-off is fully desirable in the fast-evolving commodity cloud.
Antoine Kaufmann, Tim Stamler, Simon Peter 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Thomas E. Anderson
EuroSys3
2019 Offloading distributed applications onto smartNICs using iPipe
abstract
Emerging Multicore SoC SmartNICs, enclosing rich computing resources (e.g., a multicore processor, onboard DRAM, accelerators, programmable DMA engines), hold the potential to offload generic datacenter server tasks. However, it is unclear how to use a SmartNIC efficiently and maximize the offloading benefits, especially for distributed applications. Towards this end, we characterize four commodity SmartNICs and summarize the offloading performance implications from four perspectives: traffic control, computing capability, onboard memory, and host communication.
Ming Liu 0027, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter 0001
SIGCOMM5
2019 E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers
Ming Liu 0027, Simon Peter 0001, Arvind Krishnamurthy, Phitchaya Mangpo Phothilimthana
USENIX ATC2
2018 Floem: A Programming System for NIC-Accelerated Network Applications
Phitchaya Mangpo Phothilimthana, Ming Liu 0027, Antoine Kaufmann, Simon Peter 0001, Rastislav Bodík, Thomas E. Anderson
OSDI4
2017 Evaluating the Power of Flexible Packet Processing for Network Resource Allocation
Naveen Kr. Sharma, Antoine Kaufmann, Thomas E. Anderson, Arvind Krishnamurthy, Jacob Nelson 0001, Simon Peter 0001
NSDI6
2017 Strata: A Cross Media File System
abstract
Current hardware and application storage trends put immense pressure on the operating system's storage subsystem. On the hardware side, the market for storage devices has diversified to a multi-layer storage topology spanning multiple orders of magnitude in cost and performance. Above the file system, applications increasingly need to process small, random IO on vast data sets with low latency, high throughput, and simple crash consistency. File systems designed for a single storage layer cannot support all of these demands together.
Youngjin Kwon, Henrique Fingler, Tyler Hunt, Simon Peter 0001, Emmett Witchel, Thomas E. Anderson
SOSP4
2017 Ryoan: A Distributed Sandbox for Untrusted Computation on Secret Data
abstract
Users of modern data-processing services such as tax preparation or genomic screening are forced to trust them with data that the users wish to keep secret. Ryoan 1 protects secret data while it is processed by services that the data owner does not trust. Accomplishing this goal in a distributed setting is difficult, because the user has no control over the service providers or the computational platform. Confining code to prevent it from leaking secrets is notoriously difficult, but Ryoan benefits from new hardware and a request-oriented data model. Ryoan provides a distributed sandbox, leveraging hardware enclaves (e.g., Intel’s software guard extensions (SGX) [40]) to protect sandbox instances from potentially malicious computing platforms. The protected sandbox instances confine untrusted data-processing modules to prevent leakage of the user’s input data. Ryoan is designed for a request-oriented data model, where confined modules only process input once and do not persist state about the input. We present the design and prototype implementation of Ryoan and evaluate it on a series of challenging problems including email filtering, health analysis, image processing and machine translation.
Tyler Hunt, Zhiting Zhu, Yuanzhong Xu, Simon Peter 0001, Emmett Witchel
ACM Trans. Comput. Syst.4
2016 High Performance Packet Processing with FlexNIC
abstract
The recent surge of network I/O performance has put enormous pressure on memory and software I/O processing sub systems. We argue that the primary reason for high memory and processing overheads is the inefficient use of these resources by current commodity network interface cards (NICs). We propose FlexNIC, a flexible network DMA interface that can be used by operating systems and applications alike to reduce packet processing overheads. FlexNIC allows services to install packet processing rules into the NIC, which then executes simple operations on packets while exchanging them with host memory. Thus, our proposal moves some of the packet processing traditionally done in software to the NIC, where it can be done flexibly and at high speed.
Antoine Kaufmann, Simon Peter 0001, Naveen Kr. Sharma, Thomas E. Anderson, Arvind Krishnamurthy
ASPLOS2
2016 Ryoan: A Distributed Sandbox for Untrusted Computation on Secret Data
Tyler Hunt, Zhiting Zhu, Yuanzhong Xu, Simon Peter 0001, Emmett Witchel
OSDI4
2016 Coordinated and Efficient Huge Page Management with Ingens
Youngjin Kwon, Hangchen Yu, Simon Peter 0001, Christopher J. Rossbach, Emmett Witchel
OSDI3
2016 Sweet Spots and Limits for Virtualization
abstract
This year at VEE, we added a panel to discuss the state of virtualization: what problems are solved? what problems are important? and what problems may not be worth solving? The panelist are experts in areas ranging from hardware virtualization up to language-level virtualization.
Carl A. Waldspurger, Emery D. Berger, Abhishek Bhattacharjee, Kevin T. Pedretti, Simon Peter 0001, Christopher J. Rossbach
VEE5
2016 Arrakis: The Operating System Is the Control Plane
abstract
Recent device hardware trends enable a new approach to the design of network server operating systems. In a traditional operating system, the kernel mediates access to device hardware by server applications to enforce process isolation as well as network and disk security. We have designed and implemented a new operating system, Arrakis, that splits the traditional role of the kernel in two. Applications have direct access to virtualized I/O devices, allowing most I/O operations to skip the kernel entirely, while the kernel is re-engineered to provide network and disk protection without kernel mediation of every operation. We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2 to 5 × in latency and 9 × throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation.
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
ACM Trans. Comput. Syst.1
2015 Subways: a case for redundant, inexpensive data center edge links
abstract
As network demand increases, data center network operators face a number of challenges including the need to add capacity to the network. Unfortunately, network upgrades can be an expensive proposition, particularly at the edge of the network where most of the network's cost lies.
Vincent Liu 0001, Danyang Zhuo, Simon Peter 0001, Arvind Krishnamurthy, Thomas E. Anderson
CoNEXT3
2015 FlexNIC: Rethinking Network DMA
Antoine Kaufmann, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy
HotOS2
2014 Towards High-Performance Application-Level Storage Management
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Thomas E. Anderson, Arvind Krishnamurthy, Mark Zbikowski, Doug Woos
HotStorage1
2014 Arrakis: The Operating System is the Control Plane
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe
OSDI1
2014 One tunnel is (often) enough
abstract
A longstanding problem with the Internet is that it is vulnerable to outages, black holes, hijacking and denial of service. Although architectural solutions have been proposed to address many of these issues, they have had difficulty being adopted due to the need for widespread adoption before most users would see any benefit. This is especially relevant as the Internet is increasingly used for applications where correct and continuous operation is essential.
Simon Peter 0001, Umar Javed, Qiao Zhang 0001, Doug Woos, Thomas E. Anderson, Arvind Krishnamurthy
SIGCOMM1
2013 Arrakis: A Case for the End of the Empire
Simon Peter 0001, Thomas E. Anderson
HotOS1
2013 Expressive privacy control with pseudonyms
abstract
As personal information increases in value, the incentives for remote services to collect as much of it as possible increase as well. In the current Internet, the default assumption is that all behavior can be correlated using a variety of identifying information, not the least of which is a user's IP address. Tools like Tor, Privoxy, and even NATs, are located at the opposite end of the spectrum and prevent any behavior from being linked. Instead, our goal is to provide users with more control over linkability---which activites of the user can be correlated at the remote services---not necessarily more anonymity.
Seungyeop Han, Vincent Liu 0001, Qifan Pu, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy, David Wetherall
SIGCOMM4
2012 Fay: Extensible Distributed Tracing from Kernels to Clusters
abstract
Fay is a flexible platform for the efficient collection, processing, and analysis of software execution traces. Fay provides dynamic tracing through use of runtime instrumentation and distributed aggregation within machines and across clusters. At the lowest level, Fay can be safely extended with new tracing primitives, including even untrusted, fully optimized machine code, and Fay can be applied to running user-mode or kernel-mode software without compromising system stability. At the highest level, Fay provides a unified, declarative means of specifying what events to trace, as well as the aggregation, processing, and analysis of those events. We have implemented the Fay tracing platform for Windows and integrated it with two powerful, expressive systems for distributed programming. Our implementation is easy to use, can be applied to unmodified production systems, and provides primitives that allow the overhead of tracing to be greatly reduced, compared to previous dynamic tracing platforms. To show the generality of Fay tracing, we reimplement, in experiments, a range of tracing strategies and several custom mechanisms from existing tracing frameworks. Fay shows that modern techniques for high-level querying and data-parallel processing of disagreggated data streams are well suited to comprehensive monitoring of software execution in distributed systems. Revisiting a lesson from the late 1960s [Deutsch and Grant 1971], Fay also demonstrates the efficiency and extensibility benefits of using safe, statically verified machine code as the basis for low-level execution tracing. Finally, Fay establishes that, by automatically deriving optimized query plans and code for safe extensions, the expressiveness and performance of high-level tracing queries can equal or even surpass that of specialized monitoring tools.
Úlfar Erlingsson, Marcus Peinado, Simon Peter 0001, Mihai Budiu, Gloria Mainar-Ruiz
ACM Trans. Comput. Syst.3
2012 A Declarative Language Approach to Device Configuration
abstract
C remains the language of choice for hardware programming (device drivers, bus configuration, etc.): it is fast, allows low-level access, and is trusted by OS developers. However, the algorithms required to configure and reconfigure hardware devices and interconnects are becoming more complex and diverse, with the added burden of legacy support, “quirks,” and hardware bugs to work around. Even programming PCI bridges in a modern PC is a surprisingly complex problem, and is getting worse as new functionality such as hotplug appears. Existing approaches use relatively simple algorithms, hard-coded in C and closely coupled with low-level register access code, generally leading to suboptimal configurations. We investigate the merits and drawbacks of a new approach: separating hardware configuration logic (algorithms to determine configuration parameter values) from mechanism (programming device registers). The latter we keep in C, and the former we encode in a declarative programming language with constraint-satisfaction extensions. As a test case, we have implemented full PCI configuration, resource allocation, and interrupt assignment in the Barrelfish research operating system, using a concise expression of efficient algorithms in constraint logic programming. We show that the approach is tractable, and can successfully configure a wide range of PCs with competitive runtime cost. Moreover, it requires about half the code of the C-based approach in Linux while offering considerably more functionality. Additionally it easily accommodates adaptations such as hotplug, fixed regions, and “quirks.”
Adrian Schüpbach, Andrew Baumann, Timothy Roscoe, Simon Peter 0001
ACM Trans. Comput. Syst.4
2011 A declarative language approach to device configuration
abstract
C remains the language of choice for hardware programming (device drivers, bus configuration, etc.): it is fast, allows low-level access, and is trusted by OS developers. However, the algorithms required to configure and reconfigure hardware devices and interconnects are becoming more complex and diverse, with the added burden of legacy support, quirks, and hardware bugs to work around. Even programming PCI bridges in a modern PC is a surprisingly complex problem, and is getting worse as new functionality such as hotplug appears. Existing approaches use relatively simple algorithms, hard-coded in C and closely coupled with low-level register access code, generally leading to suboptimal configurations.
Adrian Schüpbach, Andrew Baumann, Timothy Roscoe, Simon Peter 0001
ASPLOS4
2011 Fay: extensible distributed tracing from kernels to clusters
abstract
Fay is a flexible platform for the efficient collection, processing, and analysis of software execution traces. Fay provides dynamic tracing through use of runtime instrumentation and distributed aggregation within machines and across clusters. At the lowest level, Fay can be safely extended with new tracing primitives, including even untrusted, fully-optimized machine code, and Fay can be applied to running user-mode or kernel-mode software without compromising system stability. At the highest level, Fay provides a unified, declarative means of specifying what events to trace, as well as the aggregation, processing, and analysis of those events.
Úlfar Erlingsson, Marcus Peinado, Simon Peter 0001, Mihai Budiu
SOSP3
2009 Your computer is already a distributed system. Why isn't your OS?
Andrew Baumann, Simon Peter 0001, Adrian Schüpbach, Akhilesh Singhania, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs
HotOS2
2009 The multikernel: a new OS architecture for scalable multicore systems
abstract
Commodity computer systems contain more and more processor cores and exhibit increasingly diverse architectural tradeoffs, including memory hierarchies, interconnects, instruction sets and variants, and IO configurations. Previous high-performance computing systems have scaled in specific cases, but the dynamic nature of modern client and server workloads, coupled with the impossibility of statically optimizing an OS for all workloads and hardware variants pose serious challenges for operating system structures.
Andrew Baumann, Paul Barham 0001, Pierre-Évariste Dagand, Tim Harris 0001, Rebecca Isaacs, Simon Peter 0001, Timothy Roscoe, Adrian Schüpbach, Akhilesh Singhania
SOSP6
2008 30 seconds is not enough!: a study of operating system timer usage
abstract
The basic system timer facilities used by applications and OS kernels for scheduling timeouts and periodic activities have remained largely unchanged for decades, while hardware architectures and application loads have changed radically. This raises concerns with CPU overhead power management and application responsiveness.
Simon Peter 0001, Andrew Baumann, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs
EuroSys1