Marios Kogias

dblp:144/7733 · also Marios-Evangelos Kogias · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0006-7034-5284ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 7 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CheckWork: Enabling Trace-Driven Analysis of Checkpointing Overhead in Distributed ML Training
abstract
Checkpointing is a fundamental mechanism for fault tolerance in large-scale distributed machine learning (ML) training, but it can introduce significant overhead due to interactions with compute and communication phases. Evaluating checkpointing strategies on production clusters is costly, difficult to reproduce, and limits the ability to systematically explore alternative designs. We present CheckWork, a framework that generates checkpoint-aware Chakra execution traces by augmenting training DAGs with checkpoint operations. These traces enable realistic simulation of checkpointing strategies using existing system simulators. Our evaluation shows that the generated traces reproduce checkpointing behaviour previously observed on real test-beds and reported in systems such as CheckFreq and Gemini, capturing both computational overheads and network-level communication interference. Using the traces, we also analyse interference between checkpoint traffic and training communication in multi-job environments and propose a mechanism to mitigate this effect.
Eldar Hasanov, Adel Sefiane, Alireza Farshin, Marios Kogias
APNet4
2026 CAEC: Confidential, Attestable, and Efficient Inter-CVM Communication with Arm CCA
abstract
Confidential Virtual Machines (CVMs) are increasingly adopted to protect sensitive workloads from privileged adversaries such as the hypervisor. While they provide strong isolation guarantees, existing CVM architectures lack first-class mechanisms for inter-CVM data sharing due to their disjoint memory model, making inter-CVM data exchange a performance bottleneck in compartmentalized or collaborative multi-CVM systems. Under this model, a CVM's accessible memory is either shared with the hypervisor or protected from both the hypervisor and all other CVMs. This design simplifies reasoning about memory ownership; however, it fundamentally precludes plaintext data sharing between CVMs because all inter-CVM communication must pass through hypervisor-accessible memory, requiring costly encryption and decryption to preserve confidentiality and integrity. In this paper, we introduce CAEC, a system that enables protected memory sharing between CVMs. CAEC builds on Arm Confidential Compute Architecture (CCA) and extends its firmware to support Confidential Shared Memory (CSM), a memory region securely shared between multiple CVMs while remaining inaccessible to the hypervisor and all non-participating CVMs. CAEC's design is fully compatible with CCA hardware and introduces only a modest increase (6%) in CCA firmware code size. CAEC delivers substantial performance benefits across a range of workloads. For instance, inter-CVM communication over CAEC achieves up to 209x reduction in CPU cycles compared to encryption-based mechanisms over hypervisor-accessible shared memory. By combining high performance, strong isolation guarantees, and attestable sharing semantics, CAEC provides a practical and scalable foundation for the next generation of trusted multi-CVM services across both edge and cloud environments.
Sina Abdollahi, Amir Al Sadi, David Kotz, Marios Kogias, Hamed Haddadi 0001
EuroS&P4
2026 Tyche: Composable Isolation as a Foundation to Manage Trust in the Cloud
abstract
Cloud workloads combine software components from different parties to process sensitive data. Each component has its own trust model - it must protect its assets from the rest of the system, yet share sensitive data with components it cannot trust to keep confidential. This tension requires composing isolation boundaries for confidentiality and encapsulation. Unfortunately, the cloud offers no direct way to compose such boundaries, forcing tenants to assemble, deploy, and maintain their own solutions. This paper shifts that burden back to the infrastructure by making composable, attestable isolation a first-class systems abstraction. We present Tyche, a security monitor that centers isolation around a unified composable abstraction: security domains (SDs). An SD is an execution environment whose access to machine resources - memory, cores, devices - is controlled through explicit capabilities. A small set of capability operations enables SDs to partition, share, and reclaim resources; by nesting recursively, SDs compose attestable trust boundaries for confidentiality and encapsulation. Tyche attests these compositions, providing end-to-end security guarantees for workloads made of mutually distrustful components. As a first-class cloud primitive, this single abstraction subsumes enclaves, sandboxes, CVMs, and their compositions. Tyche provides composable isolation without sacrificing compatibility with existing hardware and software stacks. It runs on commodity x86 64 hardware without security extensions, and a RISC-V prototype demonstrates portability across platforms. Our SDK composes isolation for unmodified workloads within SDs with minimal overhead. In a confidential LLM inference scenario with mutually distrustful users, model owners, and cloud providers, the slowdown is just 2% compared to bare-metal Linux.
Adrien Ghosn, Charly Castes, Neelu S. Kalani, Yuchen Qian, Marios Kogias, Edouard Bugnion
EuroS&P5
2026 KRAKENGUARD: Towards Fine-Grained eBPF Isolation
Jainil Patel, Lucas Graeff Buhl-Nielsen, Adrien Ghosn, Marios Kogias
NSDI4
2025 SIRD: A Sender-Informed, Receiver-Driven Datacenter Transport Protocol
Konstantinos Prasopoulos, Ryan Kosta, Edouard Bugnion, Marios Kogias
NSDI4
2025 DORADD: Deterministic Parallel Execution in the Era of Microsecond-Scale Computing
abstract
Deterministic parallelism is a key building block for distributed and fault-tolerant systems that offers substantial performance benefits while guaranteeing determinism. By studying existing deterministically parallel systems (DPS), we identify certain design pitfalls, such as batched execution and inefficient runtime synchronization, that preclude them from meeting the demands of μs-scale and high-throughput distributed systems deployed in modern datacenters.
Zhengqing Liu, Musa Unal, Matthew J. Parkinson, Marios Kogias
PPoPP4
2023 Creating Trust by Abolishing Hierarchies
abstract
Software is going through a trust crisis. Privileged code is no longer trusted and processes insufficiently protect user code from unverified libraries. While usually treated separately, confidential computing and program compartmentalization are both symptoms of the same problem, deeply rooted in hierarchical commodity systems: privileged software's monopoly over isolation.
Charly Castes, Adrien Ghosn, Neelu S. Kalani, Yuchen Qian, Marios Kogias, Mathias Payer, Edouard Bugnion
HotOS5
2023 Towards (Really) Safe and Fast Confidential I/O
abstract
Confidential cloud computing enables cloud tenants to distrust their service provider. Achieving confidential computing solutions that provide concrete security guarantees requires not only strong mechanisms, but also carefully designed software interfaces. In this paper, we make the observation that confidential I/O interfaces, caught in the tug-of-war between performance and security, fail to address both at a time when confronted to interface vulnerabilities and observability by the untrusted host. We discuss the problem of safe I/O interfaces in confidential computing, its implications and challenges, and devise research paths to achieve confidential I/O interfaces that are both safe and fast.
Hugo Lefeuvre, David Chisnall, Marios Kogias, Pierre Olivier
HotOS3
2023 Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal Scheduling
abstract
Datacenter applications expect microsecond-scale service times and tightly bound tail latency, with future workloads expected to be even more demanding. To address this challenge, state-of-the-art runtimes employ theoretically optimal scheduling policies, namely a single request queue and strict preemption.
Rishabh Iyer 0002, Musa Unal, Marios Kogias, George Candea
SOSP3
2023 When Concurrency Matters: Behaviour-Oriented Concurrency
abstract
Expressing parallelism and coordination is central for modern concurrent programming. Many mechanisms exist for expressing both parallelism and coordination. However, the design decisions for these two mechanisms are tightly intertwined. We believe that the interdependence of these two mechanisms should be recognised and achieved through a single, powerful primitive. We are not the first to realise this: the prime example is actor model programming, where parallelism arises through fine-grained decomposition of a program’s state into actors that are able to execute independently in parallel. However, actor model programming has a serious pain point: updating multiple actors as a single atomic operation is a challenging task. We address this pain point by introducing a new concurrency paradigm: Behaviour-Oriented Concurrency (BoC). In BoC, we are revisiting the fundamental concept of a behaviour to provide a more transactional concurrency model. BoC enables asynchronously creating atomic and ordered units of work with exclusive access to a collection of independent resources. In this paper, we describe BoC informally in terms of examples, which demonstrate the advantages of exclusive access to several independent resources, as well as the need for ordering. We define it through a formal model. We demonstrate its practicality by implementing a C++ runtime. We argue its applicability through the Savina benchmark suite: benchmarks in this suite can be more compactly represented using BoC in place of Actors, and we observe comparable, if not better, performance.
Luke Cheeseman, Matthew J. Parkinson, Sylvan Clebsch, Marios Kogias, Sophia Drossopoulou, David Chisnall, Tobias Wrigstad, Paul Liétar
Proc. ACM Program. Lang.4
2021 Enclosure: language-based restriction of untrusted libraries
abstract
Programming languages and systems have failed to address the security implications of the increasingly frequent use of public libraries to construct modern software. Most languages provide tools and online repositories to publish, import, and use libraries; however, this double-edged sword can incorporate a large quantity of unknown, unchecked, and unverified code into an application. The risk is real, as demonstrated by malevolent actors who have repeatedly inserted malware into popular open-source libraries.
Adrien Ghosn, Marios Kogias, Mathias Payer, James R. Larus, Edouard Bugnion
ASPLOS2
2021 Benchmarking, analysis, and optimization of serverless function snapshots
abstract
Serverless computing has seen rapid adoption due to its high scalability and flexible, pay-as-you-go billing model. In serverless, developers structure their services as a collection of functions, sporadically invoked by various events like clicks. High inter-arrival time variability of function invocations motivates the providers to start new function instances upon each invocation, leading to significant cold-start delays that degrade user experience. To reduce cold-start latency, the industry has turned to snapshotting, whereby an image of a fully-booted function is stored on disk, enabling a faster invocation compared to booting a function from scratch.
Dmitrii Ustiugov, Plamen Petrov 0003, Marios Kogias, Edouard Bugnion, Boris Grot
ASPLOS3
2021 SmartHarvest: harvesting idle CPUs safely and efficiently in the cloud
abstract
We can increase the efficiency of public cloud datacenters by harvesting allocated but temporarily idling CPU cores from customer virtual machines (VMs) to run batch or analytics workloads. Even small efficiency gains translate into substantial savings, since provisioning and operating a datacenter costs hundreds of millions of dollars per year. The main challenge is to harvest idle cores with little or no impact on customer VMs, which could be running latency-sensitive services and are essentially black-boxes to the cloud provider.
Kapil Arya, Marios Kogias, Manohar Vanga, Aditya Bhandari, Neeraja J. Yadwadkar, Siddhartha Sen 0001, Sameh Elnikety, Christoforos E. Kozyrakis, Ricardo Bianchini
EuroSys3
2021 When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with Perséphone
abstract
This paper introduces Perséphone, a kernel-bypass OS scheduler designed to minimize tail latency for applications executing at microsecond-scale and exhibiting wide service time distributions. Perséphone integrates a new scheduling policy, Dynamic Application-aware Reserved Cores (DARC), that reserves cores for requests with short processing times. Unlike existing kernel-bypass schedulers, DARC is not work conserving. DARC profiles application requests and leaves a small number of cores idle when no short requests are in the queue, so when short requests do arrive, they are not blocked by longer-running ones. Counter-intuitively, leaving cores idle lets DARC maintain lower tail latencies at higher utilization, reducing the overall number of cores needed to serve the same workloads and consequently better utilizing the datacenter resources.
Henri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias, Boon Thau Loo, Linh T. X. Phan, Irene Zhang
SOSP4
2020 Tail-tolerance as a Systems Principle not a Metric
abstract
Tail-latency tolerance (or just simply tail-tolerance) is the ability for a system to deliver a response with low-latency nearly all the time. It it typically expressed as a system metric (e.g., the 99th or 99.99th percentile latency) or as a service-level objective (e.g., the maximum throughput so that the tail latency is below a desired threshold).
Marios Kogias, Edouard Bugnion
APNet1
2020 Bypassing the load balancer without regrets
abstract
Load balancers are a ubiquitous component of cloud deployments and the cornerstone of workload elasticity. Load balancers can significantly affect the end-to-end application latency with their load balancing decisions, and constitute a significant portion of cloud tenant expenses.
Marios Kogias, Rishabh Iyer 0002, Edouard Bugnion
SoCC1
2020 HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter services
abstract
Cloud platform services must simultaneously be scalable, meet low tail latency service-level objectives, and be resilient to a combination of software, hardware, and network failures. Replication plays a fundamental role in meeting both the scalability and the fault-tolerance requirement, but is subject to opposing requirements: (1) scalability is typically achieved by relaxing consistency; (2) fault-tolerance is typically achieved through the consistent replication of state machines. Adding nodes to a system can therefore either increase performance at the expense of consistency, or increase resiliency at the expense of performance.
Marios Kogias, Edouard Bugnion
EuroSys1
2019 Lancet: A self-correcting Latency Measuring Tool
Marios Kogias, Stephen Mallon, Edouard Bugnion
USENIX ATC1
2019 R2P2: Making RPCs first-class datacenter citizens
Marios Kogias, George Prekas, Adrien Ghosn, Jonas Fietz, Edouard Bugnion
USENIX ATC1
2017 ZygOS: Achieving Low Tail Latency for Microsecond-scale Networked Tasks
abstract
This paper focuses on the efficient scheduling on multicore systems of very fine-grain networked tasks, which are the typical building block of online data-intensive applications. The explicit goal is to deliver high throughput (millions of remote procedure calls per second) for tail latency service-level objectives that are a small multiple of the task size.
George Prekas, Marios Kogias, Edouard Bugnion
SOSP2