Kostis Kaffes

dblp:237/0852 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-0517-7206ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Enabling Fast Networking in the Public Cloud
abstract
Despite a decade of research, most high-performance userspace network stacks remain impractical for public cloud tenants developing their applications atop Virtual Machines (VMs). We identify two root causes: (1) reliance on specialized NIC features (e.g., flow steering, deep buffers) absent in commodity cloud vNICs, and (2) rigid execution models ill-suited to diverse application needs. We present Machnet, a highperformance and flexible userspace network stack designed for public cloud VMs. Machnet uses only a minimal set of vNIC features that any major cloud provider supports. It also relies on a microkernel architecture to enable flexible application execution. We evaluate Machnet across three major public clouds and on production-grade applications, including a key-value store, an HTTP server, and a statemachine replication system. We release Machnet at https: //github.com/microsoft/machnet.
Alireza Sanaee, Vahab Jabrayilov, Ilias Marinos, Farbod Shahinfar, Divyanshu Saxena, Gianni Antichi, Kostis Kaffes
ASPLOS (2)7
2026 Critical Path Guided Decision Making with CALLIGATOR
Meghna Pancholi, Lee Baugh, Olaf Schnapauff, David E. Culler, Kostis Kaffes, Yu Gan 0002, Brent E. Stephens
SIGCOMM5
2025 Wave: Offloading Resource Management to SmartNIC Cores
abstract
SmartNICs are increasingly deployed in datacenters to offload tasks from server CPUs, improving the efficiency and flexibility of datacenter security, networking and storage. Optimizing cloud server efficiency in this way is critically important to ensure that virtually all server resources are available to paying customers. Userspace system software, specifically, decision-making tasks performed by various operating system subsystems, is particularly well suited for execution on mid-tier SmartNIC ARM cores. To this end, we introduce Wave, a framework for offloading userspace system software to processes/agents running on the SmartNIC. Wave uses Linux userspace systems to better align system functionality with SmartNIC capabilities. It also introduces a new host-SmartNIC communication API that enables offloading of even μs-scale system software. To evaluate Wave, we offloaded preexisting userspace system software including kernel thread scheduling, memory management, and an RPC stack to SmartNIC ARM cores, which showed a performance degradation of 1.1%-7.4% in an apples-to-apples comparison with on-host implementations. Wave recovered host resources consumed by on-host system software for memory management (saving 16 host cores), RPCs (saving 8 host cores), and virtual machines (an 11.2% performance improvement). Wave highlights the potential for rethinking system software placement in modern datacenters, unlocking new opportunities for efficiency and scalability.
Jack Tigar Humphries, Neel Natu, Kostis Kaffes, Stanko Novakovic, Henry M. Levy, David E. Culler, Christoforos E. Kozyrakis
ASPLOS (3)3
2025 ALAP: Intent-Based Serverless Computing via Delayed Decision-Making
abstract
The lack of an interface that allows users to express their intent, to trade off latency versus cost, continues to be a hurdle in the adoption of serverless computing for latency-critical and cost-constrained applications. Existing systems shy away from providing a user-intent knob mainly because they make resource allocation decisions as soon as a function is registered, when inputs to the function are mostly unavailable. We find that function performance and resource utilization can vary greatly (up to 6× in latency and 4.78× in utilization) across inputs. Thus, to enable intent-based serverless computing, our key insight is to delay making resource allocation decisions until after function inputs are available so that resources can be allocated independently for each input and resource type. We introduce ALAP, a resource management framework for serverless systems that makes decisions As Late As Possible to right-size each invocation and meet user-intent efficiently. ALAP uses an online learning agent to predict the required amount of resources to meet an invocation's constraints and schedules these right-sized containers in a cold-start-aware manner. For a range of functions and inputs, ALAP adapts to user intents: it reduces latency and cost by up to 17.5× (3× on avg.) and 13× (3.7× on avg.), respectively, and latency and cost SLO violations by 1.3-2.3× and 1.6-2.2×, respectively, while nearly eliminating CPU and memory waste compared to four state-of-the-art systems.
Prasoon Sinha, Kostis Kaffes, Neeraja J. Yadwadkar
SoCC2
2024 Rethinking the Networking Stack for Serverless Environments: A Sidecar Approach
abstract
Serverless platforms rely on legacy networking stacks for communication and data movement. We quantitatively analyze the performance of these stacks and show their mismatch with highly consolidated, virtualized modern serverless environments, focusing on Firecracker, the most common serverless virtualization framework. As serverless applications grow in complexity and interaction, the resulting network bottleneck is a prime source of user-perceived, end-to-end latency. In this paper, we present a detailed vision of a new, sidecar-based networking stack for serverless environments. Our primary design goal is to provide low-overhead networking while maintaining existing security guarantees. We outline the research challenges in both the control and the data plane that the community needs to tackle before such a sidecar architecture can be used in practice.
Vishwanath Seshagiri, Vahab Jabrayilov, Avani Wildani, Kostis Kaffes
SoCC5
2024 LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space Scheduling
abstract
Modern cloud applications are prone to high tail latencies since their requests typically follow highly-dispersive distributions. Prior work has proposed both OS- and systemlevel solutions to reduce tail latencies for microsecond-scale workloads through better scheduling. Unfortunately, existing approaches like customized dataplane OSes, require significant OS changes, experience scalability limitations, or do not reach the full performance capabilities hardware offers. We propose LibPreemptible, a preemptive user-level threading library that is flexible, lightweight, and scalable. LibPreemptible is based on three key techniques: 1) a fast and lightweight hardware mechanism for delivery of timed interrupts, 2) a general-purpose user-level scheduling interface, and 3) an API for users to express adaptive scheduling policies tailored to the needs of their applications. Compared to the prior state-of-the-art scheduling system Shinjuku, our system achieves significant tail latency and throughput improvements for various workloads without the need to modify the kernel. We also demonstrate the flexibility of LibPreemptible across scheduling policies for real applications experiencing varying load levels and characteristics.
Nikita Lazarev, David Koufaty, Tenny Yin, Andy Anderson, Zhiru Zhang, G. Edward Suh, Kostis Kaffes, Christina Delimitrou
HPCA8
2022 A Progress Report on DBOS: A Database-oriented Operating System
Qian Li 0027, Peter Kraft, Kostis Kaffes, Athinagoras Skiadopoulos, Deeptaanshu Kumar, Michael J. Cafarella, Goetz Graefe, Jeremy Kepner, Christoforos E. Kozyrakis, Michael Stonebraker, Lalith Suresh 0001, Matei Zaharia
CIDR3
2022 Hermod: principled and practical scheduling for serverless functions
abstract
Serverless computing has seen rapid growth due to the ease-of-use and cost-efficiency it provides. However, function scheduling, a critical component of serverless systems, has been overlooked. In this paper, we take a fist-principles approach toward designing a scheduler that caters to the unique characteristics of serverless functions as seen in real-world deployments. We first create a taxonomy of scheduling policies along three dimensions. Next, we use simulation to explore the scheduling policy space and show that frequently used features such as late binding and random load balancing are sub-optimal for common execution time distributions and load ranges. We use these insights to design Hermod, a scheduler for serverless functions with two key characteristics. First, to avoid head-of-line blocking due to high function execution time variability, Hermod uses a combination of early binding and processor sharing for scheduling at individual worker machines. Second, Hermod is cost, load, and locality-aware. It improves consolidation at low load, it employs least-loaded balancing at high load to retain high performance, and it reduces the number of cold starts compared to pure load-based policies. We implement Hermod for Apache OpenWhisk and demonstrate that, for the case of the function patterns observed in real-world traces, it achieves up to 85% lower function slowdown and 60% higher throughput compared to existing production and state-of-the-art research schedulers.
Kostis Kaffes, Neeraja J. Yadwadkar, Christoforos E. Kozyrakis
SoCC1
2021 A case against (most) context switches
abstract
Multiplexing software threads onto hardware threads and serving interrupts, VM-exits, and system calls require frequent context switches, causing high overheads and significant kernel and application complexity. We argue that context switching is an idea whose time has come and gone, and propose eliminating it through a radically different hardware threading model targeted to solve software rather than hardware problems. The new model adds a large number of hardware threads to each physical core - making thread multiplexing unnecessary - and lets software manage them. The only state change directly triggered in hardware by system calls, exceptions, and asynchronous hardware events will be blocking and unblocking hardware threads. We also present ISA extensions to allow kernel and user software to exploit this new threading model. Developers can use these extensions to eliminate interrupts and implement fast I/O without polling, exception-less system and hypervisor calls, practical microkernels, simple distributed programming models, and untrusted but fast hypervisors. Finally, we suggest practical hardware implementations and discuss the hardware and software challenges toward realizing this novel approach.
Jack Tigar Humphries, Kostis Kaffes, David Mazières, Christoforos E. Kozyrakis
HotOS2
2021 Syrup: User-Defined Scheduling Across the Stack
abstract
Suboptimal scheduling decisions in operating systems, networking stacks, and application runtimes are often responsible for poor application performance, including higher latency and lower throughput. These poor decisions stem from a lack of insight into the applications and requests the scheduler is handling and a lack of coherence and coordination between the various layers of the stack, including NICs, kernels, and applications.
Kostis Kaffes, Jack Tigar Humphries, David Mazières, Christoforos E. Kozyrakis
SOSP1
2021 DBOS: A DBMS-oriented Operating System
abstract
This paper lays out the rationale for building a completely new operating system (OS) stack. Rather than build on a single node OS together with separate cluster schedulers, distributed filesystems, and network managers, we argue that a distributed transactional DBMS should be the basis for a scalable cluster OS. We show herein that such a database OS (DBOS) can do scheduling, file management, and inter-process communication with competitive performance to existing systems. In addition, significantly better analytics can be provided as well as a dramatic reduction in code complexity through implementing OS services as standard database queries, while implementing low-latency transactions and high availability only once.
Athinagoras Skiadopoulos, Qian Li 0027, Peter Kraft, Kostis Kaffes, Daniel Hong, Shana Mathew, David Bestor, Michael J. Cafarella, Vijay Gadepally, Goetz Graefe, Jeremy Kepner, Christoforos E. Kozyrakis, Tim Kraska, Michael Stonebraker, Lalith Suresh 0001, Matei Zaharia
Proc. VLDB Endow.4
2020 Leveraging application classes to save power in highly-utilized data centers
abstract
Data center energy consumption has become an increasingly significant contributor both to greenhouse emissions and costs. To increase utilization of individual hosts and improve efficiency, most modern data centers co-locate workloads belonging to different application classes, some being latency-sensitive (LS) and others best-effort (BE) which are more tolerant to performance variation. It is therefore necessary to design mechanisms that reduce power consumption even in the resulting high-utilization environment, while preserving LS task performance. Moreover, the abundance of different workloads and the security implications of public cloud make mechanisms that rely on extensive knowledge of workload characteristics or on application-exported metrics challenging to deploy.
Kostis Kaffes, Dragos Sbirlea, Yiyan Lin, David Lo 0003, Christoforos E. Kozyrakis
SoCC1
2020 RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers
Kostis Kaffes, Zixu Chen, Zhenming Liu, Christoforos E. Kozyrakis, Ion Stoica, Xin Jin 0008
OSDI2
2019 Centralized Core-granular Scheduling for Serverless Functions
abstract
In recent years, many applications have started using serverless computing platforms primarily due to the ease of deployment and cost efficiency they offer. However, the existing scheduling mechanisms of serverless platforms fall short in catering to the unique characteristics of such applications: burstiness, short and variable execution times, statelessness and use of a single core. Specifically, the existing mechanisms fall short in meeting the requirements generated due to the combined effect of these characteristics: scheduling at a scale of millions of function invocations per second while achieving predictable performance.
Kostis Kaffes, Neeraja J. Yadwadkar, Christoforos E. Kozyrakis
SoCC1
2019 Mind the Gap: A Case for Informed Request Scheduling at the NIC
abstract
Recent research in high-throughput networked systems has established the need for centralized and preemptive request scheduling in order to achieve good hardware utilization and low tail latency for a wide variety of workloads. However, this approach is expensive to scale as it requires an increasing number of CPU cores dedicated to scheduling. Moreover, passing every request through a scheduling core introduces latency for inter-core communication and reduces the effectiveness of data preloading and caching optimizations.
Jack Tigar Humphries, Kostis Kaffes, David Mazières, Christoforos E. Kozyrakis
HotNets2
2019 Shinjuku: Preemptive Scheduling for μsecond-scale Tail Latency
Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, Christoforos E. Kozyrakis
NSDI1