VLDB 2026 Research / reviewers in the wild / expert
Simon Peter 0001
dblp:43/3917
· DBLP profile ↗
43ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0007-4748-8524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 2 first-author · 7 since 2021Systems, architecture and hardware · 18 · 3 first-author · 7 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing the GPU Memory Bottleneck with Lossless Compression for MLabstractMachine learning (ML) training and inference often process data sets far exceeding GPU memory capacity, forcing them to rely on PCIe for on-demand tensor transfers, causing critical transfer bottlenecks. Lossy compression has been proposed to relieve bottlenecks but introduces workload-dependent accuracy loss, making it complex or even prohibitive to use in existing ML deployments. Aditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon Peter 0001 |
EuroSys | 4 |
| 2026 | Prediction-Informed Power Management for General-Purpose Compute ServersabstractThis paper presents PIP (Prediction-Informed Power), a power control framework for general-purpose compute servers. PIP introduces two key innovations: (1) a machine learning-based power model that predicts the impact of hypothetical CPU throttling actions before execution, and (2) a prediction-informed control loop that selects CPU configurations to maximize performance and power utilization based on these predictions. By leveraging finegrained runtime CPU metrics, PIP can accurately estimate counterfactual power usage, allowing the control system to align power demand with the budget more quickly. Unlike traditional reactive approaches, PIP maintains effective control under frequent budget fluctuations, achieving safe oversubscription by up to 70%. Our evaluation on diverse application workloads, none of which are included in the model's training set, shows that PIP yields up to a 3.2× speedup over a state-of-the-art feedback-based system for single-application runs, and up to a 3.4× speedup for multi-application scenarios under power constraints. Jonggyu Park, Simon Peter 0001, Thomas E. Anderson |
EuroSys | 2 |
| 2026 | PASS: A Power Adaptive Storage ServerabstractPower management has become important in data centers. Since data center workloads are often dynamic, it is common practice to conserve energy by scaling resources up or down to match the workload. And since data centers often oversubscribe the power delivery infrastructure, operators can add power capping on top of these workload-proportional systems to adjust to available power. We find that this combination leaves performance on the table and provides only a limited power control range. Instead, we argue for power-adaptive systems that attempt to make the best use of the available power budget. To illustrate this approach, we built PASS, a power-adaptive storage system. PASS considers the interactions between different system components, including software and hardware, when making its power management decisions. For example, when under the same power constraint, PASS achieves 3–25× better throughput on filebench workloads than Intel's SPDK storage stack with Google Thunderbolt, a state of the art power capping system. Dedong Xie, Theano Stavrinos, Jonggyu Park, Simon Peter 0001, Baris Kasikci, Thomas E. Anderson |
EuroSys | 4 |
| 2026 | Presto: A Match-Action TCP Stack for the Terabit EraabstractWe present Presto, the first TCP stack that delivers ASIC-class performance and energy efficiency on programmable Reconfigurable Match-Action Table (RMT) pipelines, providing flexibility while retaining standard TCP semantics and POSIX socket compatibility. The key challenge in designing Presto is reconciling TCP's complex, dependent state updates with RMT's unidirectional, lock-step execution model. To overcome this challenge, Presto introduces three novel techniques: optimistic concurrency (speculative updates validated downstream), pseudo-segment injection (circular dependency resolution without stalls), and bump-in-the-wire processing (singlepass segment handling). Together, these enable TCP retransmission, reassembly, flow, and congestion control, as a pipeline of simple match-action operations. Rajath Shashidhara, Antoine Kaufmann, Simon Peter 0001 |
SIGCOMM | 3 |
| 2025 | POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter 0001, Ramachandran Ramjee, Ashish Panwar |
ASPLOS (2) | 4 |
| 2024 | Can Storage Devices be Power Adaptive?abstractPower is becoming a scarce resource for data centers, raising the need for power adaptive system design---the ability to dynamically change power consumption---to match available power. Storage makes up an increasing fraction of total data center power consumption. As such, it holds great potential to contribute to data center power adaptivity. Dedong Xie, Theano Stavrinos, Kan Zhu, Simon Peter 0001, Baris Kasikci, Thomas E. Anderson |
HotStorage | 4 |
| 2024 | (MC)2: Lazy MemCopy at the Memory Controllerabstract(MC)2is a lazy memory copy mechanism which can be used within memcpy-like functions to significantly reduce the CPU overhead for copies that are sparsely accessed. It can also hide copy latencies by enhancing the CPU’s ability to execute them asynchronously. (MC)2,s lazy memcpy avoids copying data at the time of invocation. Instead, (MC)2tracks prospective copies. If copied data is later accessed by a CPU or the cache, (MC)2uses the tracking information to lazily execute a copy, when necessary. Placing (MC)2at the memory controller puts it at the perfect vantage point to eliminate the largest source of memcpy overhead–CPU stalls due to cache misses in the critical path–while imposing minimal overhead itself. (MC)2consists of three main components: memory controller extensions that implement a lazy memcpy operation, a new instruction exposing the lazy memcpy, and a flexible software wrapper with semantics identical to memcpy. We implement and evaluate (MC)2in the gem5 simulator using a variety of microbenchmarks and workloads, including Google’s Protobuf, where (MC)2provides a $43 \%$ speedup and Linux huge page copy-on-write faults, where (MC)2provides $250 \times$ lower latency. Aditya K. Kamath, Simon Peter 0001 |
ISCA | 2 |
| 2023 | ScaleDB: A Scalable, Asynchronous In-Memory Database
Syed Akbar Mehdi, Deukyeon Hwang, Simon Peter 0001, Lorenzo Alvisi |
OSDI | 3 |
| 2022 | RDMA is Turing complete, we just did not know it yet!
Waleed Reda, Marco Canini, Dejan Kostic, Simon Peter 0001 |
NSDI | 4 |
| 2022 | FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism
Rajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon Peter 0001 |
NSDI | 4 |
| 2022 | zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IO
Tim Stamler, Deukyeon Hwang, Amanda Raybuck, Simon Peter 0001 |
OSDI | 5 |
| 2021 | Rethinking File Mapping for Persistent Memory
Ian Neal, Gefei Zuo, Eric Shiple, Tanvir Ahmed Khan 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci |
FAST | 6 |
| 2021 | Rearchitecting Linux Storage Stack for µs Latency and High Throughput
Jae-Hyun Hwang, Midhul Vuppalapati, Simon Peter 0001, Rachit Agarwal 0001 |
OSDI | 3 |
| 2021 | LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismabstractIn multi-tenant systems, the CPU overhead of distributed file systems (DFSes) is increasingly a burden to application performance. CPU and memory interference cause degraded and unstable application and storage performance, in particular for operation latency. Recent client-local DFSes for persistent memory (PM) accelerate this trend. DFS offload to SmartNICs is a promising solution to these problems, but it is challenging to fit the complex demands of a DFS onto simple SmartNIC processors located across PCIe. Jongyul Kim 0001, Insu Jang, Waleed Reda, Jaeseong Im, Marco Canini, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Emmett Witchel |
SOSP | 8 |
| 2021 | HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMabstractHigh-capacity non-volatile memory (NVM) is a new main memory tier. Tiered DRAM+NVM servers increase total memory capacity by up to 8x, but can diminish memory bandwidth by up to 7x and inflate latency by up to 63% if not managed well. We study existing hardware and software tiered memory management systems on the recently available Intel Optane DC NVM with big data applications and find that no existing system maximizes application performance on real NVM. Amanda Raybuck, Tim Stamler, Mattan Erez, Simon Peter 0001 |
SOSP | 5 |
| 2020 | Assise: Performance and Availability via Client-local NVM in a Distributed File System
Thomas E. Anderson, Marco Canini, Jongyul Kim 0001, Dejan Kostic, Youngjin Kwon, Simon Peter 0001, Waleed Reda, Henry Schuh, Emmett Witchel |
OSDI | 6 |
| 2020 | AGAMOTTO: How Persistent is your Persistent Memory Application?
Ian Neal, Ben Reeves, Benjamin Stoler, Andrew Quinn 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci |
OSDI | 6 |
| 2019 | TAS: TCP Acceleration as an OS ServiceabstractAs datacenter network speeds rise, an increasing fraction of server CPU cycles is consumed by TCP packet processing, in particular for remote procedure calls (RPCs). To free server CPUs from this burden, various existing approaches have attempted to mitigate these overheads, by bypassing the OS kernel, customizing the TCP stack for an application, or by offloading packet processing to dedicated hardware. In doing so, these approaches trade security, agility, or generality for efficiency. Neither trade-off is fully desirable in the fast-evolving commodity cloud. Antoine Kaufmann, Tim Stamler, Simon Peter 0001, Naveen Kr. Sharma, Arvind Krishnamurthy, Thomas E. Anderson |
EuroSys | 3 |
| 2019 | Offloading distributed applications onto smartNICs using iPipeabstractEmerging Multicore SoC SmartNICs, enclosing rich computing resources (e.g., a multicore processor, onboard DRAM, accelerators, programmable DMA engines), hold the potential to offload generic datacenter server tasks. However, it is unclear how to use a SmartNIC efficiently and maximize the offloading benefits, especially for distributed applications. Towards this end, we characterize four commodity SmartNICs and summarize the offloading performance implications from four perspectives: traffic control, computing capability, onboard memory, and host communication. Ming Liu 0027, Tianyi Cui, Henry Schuh, Arvind Krishnamurthy, Simon Peter 0001 |
SIGCOMM | 5 |
| 2019 | E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers
Ming Liu 0027, Simon Peter 0001, Arvind Krishnamurthy, Phitchaya Mangpo Phothilimthana |
USENIX ATC | 2 |
| 2018 | Floem: A Programming System for NIC-Accelerated Network Applications
Phitchaya Mangpo Phothilimthana, Ming Liu 0027, Antoine Kaufmann, Simon Peter 0001, Rastislav Bodík, Thomas E. Anderson |
OSDI | 4 |
| 2017 | Evaluating the Power of Flexible Packet Processing for Network Resource Allocation
Naveen Kr. Sharma, Antoine Kaufmann, Thomas E. Anderson, Arvind Krishnamurthy, Jacob Nelson 0001, Simon Peter 0001 |
NSDI | 6 |
| 2017 | Strata: A Cross Media File SystemabstractCurrent hardware and application storage trends put immense pressure on the operating system's storage subsystem. On the hardware side, the market for storage devices has diversified to a multi-layer storage topology spanning multiple orders of magnitude in cost and performance. Above the file system, applications increasingly need to process small, random IO on vast data sets with low latency, high throughput, and simple crash consistency. File systems designed for a single storage layer cannot support all of these demands together. Youngjin Kwon, Henrique Fingler, Tyler Hunt, Simon Peter 0001, Emmett Witchel, Thomas E. Anderson |
SOSP | 4 |
| 2017 | Ryoan: A Distributed Sandbox for Untrusted Computation on Secret DataabstractUsers of modern data-processing services such as tax preparation or genomic screening are forced to trust them with data that the users wish to keep secret. Ryoan 1 protects secret data while it is processed by services that the data owner does not trust. Accomplishing this goal in a distributed setting is difficult, because the user has no control over the service providers or the computational platform. Confining code to prevent it from leaking secrets is notoriously difficult, but Ryoan benefits from new hardware and a request-oriented data model. Ryoan provides a distributed sandbox, leveraging hardware enclaves (e.g., Intel’s software guard extensions (SGX) [40]) to protect sandbox instances from potentially malicious computing platforms. The protected sandbox instances confine untrusted data-processing modules to prevent leakage of the user’s input data. Ryoan is designed for a request-oriented data model, where confined modules only process input once and do not persist state about the input. We present the design and prototype implementation of Ryoan and evaluate it on a series of challenging problems including email filtering, health analysis, image processing and machine translation. Tyler Hunt, Zhiting Zhu, Yuanzhong Xu, Simon Peter 0001, Emmett Witchel |
ACM Trans. Comput. Syst. | 4 |
| 2016 | High Performance Packet Processing with FlexNICabstractThe recent surge of network I/O performance has put enormous pressure on memory and software I/O processing sub systems. We argue that the primary reason for high memory and processing overheads is the inefficient use of these resources by current commodity network interface cards (NICs). We propose FlexNIC, a flexible network DMA interface that can be used by operating systems and applications alike to reduce packet processing overheads. FlexNIC allows services to install packet processing rules into the NIC, which then executes simple operations on packets while exchanging them with host memory. Thus, our proposal moves some of the packet processing traditionally done in software to the NIC, where it can be done flexibly and at high speed. Antoine Kaufmann, Simon Peter 0001, Naveen Kr. Sharma, Thomas E. Anderson, Arvind Krishnamurthy |
ASPLOS | 2 |
| 2016 | Ryoan: A Distributed Sandbox for Untrusted Computation on Secret Data
Tyler Hunt, Zhiting Zhu, Yuanzhong Xu, Simon Peter 0001, Emmett Witchel |
OSDI | 4 |
| 2016 | Coordinated and Efficient Huge Page Management with Ingens
Youngjin Kwon, Hangchen Yu, Simon Peter 0001, Christopher J. Rossbach, Emmett Witchel |
OSDI | 3 |
| 2016 | Sweet Spots and Limits for VirtualizationabstractThis year at VEE, we added a panel to discuss the state of virtualization: what problems are solved? what problems are important? and what problems may not be worth solving? The panelist are experts in areas ranging from hardware virtualization up to language-level virtualization. Carl A. Waldspurger, Emery D. Berger, Abhishek Bhattacharjee, Kevin T. Pedretti, Simon Peter 0001, Christopher J. Rossbach |
VEE | 5 |
| 2016 | Arrakis: The Operating System Is the Control PlaneabstractRecent device hardware trends enable a new approach to the design of network server operating systems. In a traditional operating system, the kernel mediates access to device hardware by server applications to enforce process isolation as well as network and disk security. We have designed and implemented a new operating system, Arrakis, that splits the traditional role of the kernel in two. Applications have direct access to virtualized I/O devices, allowing most I/O operations to skip the kernel entirely, while the kernel is re-engineered to provide network and disk protection without kernel mediation of every operation. We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2 to 5 × in latency and 9 × throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation. Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe |
ACM Trans. Comput. Syst. | 1 |
| 2015 | Subways: a case for redundant, inexpensive data center edge linksabstractAs network demand increases, data center network operators face a number of challenges including the need to add capacity to the network. Unfortunately, network upgrades can be an expensive proposition, particularly at the edge of the network where most of the network's cost lies. Vincent Liu 0001, Danyang Zhuo, Simon Peter 0001, Arvind Krishnamurthy, Thomas E. Anderson |
CoNEXT | 3 |
| 2015 | FlexNIC: Rethinking Network DMA
Antoine Kaufmann, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy |
HotOS | 2 |
| 2014 | Towards High-Performance Application-Level Storage Management
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Thomas E. Anderson, Arvind Krishnamurthy, Mark Zbikowski, Doug Woos |
HotStorage | 1 |
| 2014 | Arrakis: The Operating System is the Control Plane
Simon Peter 0001, Jialin Li 0001, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, Timothy Roscoe |
OSDI | 1 |
| 2014 | One tunnel is (often) enoughabstractA longstanding problem with the Internet is that it is vulnerable to outages, black holes, hijacking and denial of service. Although architectural solutions have been proposed to address many of these issues, they have had difficulty being adopted due to the need for widespread adoption before most users would see any benefit. This is especially relevant as the Internet is increasingly used for applications where correct and continuous operation is essential. Simon Peter 0001, Umar Javed, Qiao Zhang 0001, Doug Woos, Thomas E. Anderson, Arvind Krishnamurthy |
SIGCOMM | 1 |
| 2013 | Arrakis: A Case for the End of the Empire
Simon Peter 0001, Thomas E. Anderson |
HotOS | 1 |
| 2013 | Expressive privacy control with pseudonymsabstractAs personal information increases in value, the incentives for remote services to collect as much of it as possible increase as well. In the current Internet, the default assumption is that all behavior can be correlated using a variety of identifying information, not the least of which is a user's IP address. Tools like Tor, Privoxy, and even NATs, are located at the opposite end of the spectrum and prevent any behavior from being linked. Instead, our goal is to provide users with more control over linkability---which activites of the user can be correlated at the remote services---not necessarily more anonymity. Seungyeop Han, Vincent Liu 0001, Qifan Pu, Simon Peter 0001, Thomas E. Anderson, Arvind Krishnamurthy, David Wetherall |
SIGCOMM | 4 |
| 2012 | Fay: Extensible Distributed Tracing from Kernels to ClustersabstractFay is a flexible platform for the efficient collection, processing, and analysis of software execution traces. Fay provides dynamic tracing through use of runtime instrumentation and distributed aggregation within machines and across clusters. At the lowest level, Fay can be safely extended with new tracing primitives, including even untrusted, fully optimized machine code, and Fay can be applied to running user-mode or kernel-mode software without compromising system stability. At the highest level, Fay provides a unified, declarative means of specifying what events to trace, as well as the aggregation, processing, and analysis of those events. We have implemented the Fay tracing platform for Windows and integrated it with two powerful, expressive systems for distributed programming. Our implementation is easy to use, can be applied to unmodified production systems, and provides primitives that allow the overhead of tracing to be greatly reduced, compared to previous dynamic tracing platforms. To show the generality of Fay tracing, we reimplement, in experiments, a range of tracing strategies and several custom mechanisms from existing tracing frameworks. Fay shows that modern techniques for high-level querying and data-parallel processing of disagreggated data streams are well suited to comprehensive monitoring of software execution in distributed systems. Revisiting a lesson from the late 1960s [Deutsch and Grant 1971], Fay also demonstrates the efficiency and extensibility benefits of using safe, statically verified machine code as the basis for low-level execution tracing. Finally, Fay establishes that, by automatically deriving optimized query plans and code for safe extensions, the expressiveness and performance of high-level tracing queries can equal or even surpass that of specialized monitoring tools. Úlfar Erlingsson, Marcus Peinado, Simon Peter 0001, Mihai Budiu, Gloria Mainar-Ruiz |
ACM Trans. Comput. Syst. | 3 |
| 2012 | A Declarative Language Approach to Device ConfigurationabstractC remains the language of choice for hardware programming (device drivers, bus configuration, etc.): it is fast, allows low-level access, and is trusted by OS developers. However, the algorithms required to configure and reconfigure hardware devices and interconnects are becoming more complex and diverse, with the added burden of legacy support, “quirks,” and hardware bugs to work around. Even programming PCI bridges in a modern PC is a surprisingly complex problem, and is getting worse as new functionality such as hotplug appears. Existing approaches use relatively simple algorithms, hard-coded in C and closely coupled with low-level register access code, generally leading to suboptimal configurations. We investigate the merits and drawbacks of a new approach: separating hardware configuration logic (algorithms to determine configuration parameter values) from mechanism (programming device registers). The latter we keep in C, and the former we encode in a declarative programming language with constraint-satisfaction extensions. As a test case, we have implemented full PCI configuration, resource allocation, and interrupt assignment in the Barrelfish research operating system, using a concise expression of efficient algorithms in constraint logic programming. We show that the approach is tractable, and can successfully configure a wide range of PCs with competitive runtime cost. Moreover, it requires about half the code of the C-based approach in Linux while offering considerably more functionality. Additionally it easily accommodates adaptations such as hotplug, fixed regions, and “quirks.” Adrian Schüpbach, Andrew Baumann, Timothy Roscoe, Simon Peter 0001 |
ACM Trans. Comput. Syst. | 4 |
| 2011 | A declarative language approach to device configurationabstractC remains the language of choice for hardware programming (device drivers, bus configuration, etc.): it is fast, allows low-level access, and is trusted by OS developers. However, the algorithms required to configure and reconfigure hardware devices and interconnects are becoming more complex and diverse, with the added burden of legacy support, quirks, and hardware bugs to work around. Even programming PCI bridges in a modern PC is a surprisingly complex problem, and is getting worse as new functionality such as hotplug appears. Existing approaches use relatively simple algorithms, hard-coded in C and closely coupled with low-level register access code, generally leading to suboptimal configurations. Adrian Schüpbach, Andrew Baumann, Timothy Roscoe, Simon Peter 0001 |
ASPLOS | 4 |
| 2011 | Fay: extensible distributed tracing from kernels to clustersabstractFay is a flexible platform for the efficient collection, processing, and analysis of software execution traces. Fay provides dynamic tracing through use of runtime instrumentation and distributed aggregation within machines and across clusters. At the lowest level, Fay can be safely extended with new tracing primitives, including even untrusted, fully-optimized machine code, and Fay can be applied to running user-mode or kernel-mode software without compromising system stability. At the highest level, Fay provides a unified, declarative means of specifying what events to trace, as well as the aggregation, processing, and analysis of those events. Úlfar Erlingsson, Marcus Peinado, Simon Peter 0001, Mihai Budiu |
SOSP | 3 |
| 2009 | Your computer is already a distributed system. Why isn't your OS?
Andrew Baumann, Simon Peter 0001, Adrian Schüpbach, Akhilesh Singhania, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs |
HotOS | 2 |
| 2009 | The multikernel: a new OS architecture for scalable multicore systemsabstractCommodity computer systems contain more and more processor cores and exhibit increasingly diverse architectural tradeoffs, including memory hierarchies, interconnects, instruction sets and variants, and IO configurations. Previous high-performance computing systems have scaled in specific cases, but the dynamic nature of modern client and server workloads, coupled with the impossibility of statically optimizing an OS for all workloads and hardware variants pose serious challenges for operating system structures. Andrew Baumann, Paul Barham 0001, Pierre-Évariste Dagand, Tim Harris 0001, Rebecca Isaacs, Simon Peter 0001, Timothy Roscoe, Adrian Schüpbach, Akhilesh Singhania |
SOSP | 6 |
| 2008 | 30 seconds is not enough!: a study of operating system timer usageabstractThe basic system timer facilities used by applications and OS kernels for scheduling timeouts and periodic activities have remained largely unchanged for decades, while hardware architectures and application loads have changed radically. This raises concerns with CPU overhead power management and application responsiveness. Simon Peter 0001, Andrew Baumann, Timothy Roscoe, Paul Barham 0001, Rebecca Isaacs |
EuroSys | 1 |