EDBT 2026 Demo / reviewers in the wild / expert
David Kozhaya
dblp:170/0240
· DBLP profile ↗
15ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-5453-656XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Security and privacy · 3 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pallas and Aegis: Rollback Resilience in TEE-Aided Blockchain Consensus
Jeremie Decouchant, David Kozhaya, Vincent Rahli, Jiangshan Yu |
NDSS | 2 |
| 2025 | On Real-Time Guarantees in Intel SGX and TDXabstractTrusted execution environments (TEE) represent a major technological breakthrough that provide strong confidentiality and integrity guarantees for code and data running on potentially vulnerable or untrustworthy computing systems, such as cloud, edge, embedded, mobile, or even blockchain systems. However, the performance overhead associated with TEEs still poses a limitation on the extent to which real-time (RT) sensitive applications can benefit from this technology, e.g., to run on untrusted third-party infrastructures. This work investigates various TEE-based architectures spanning from process-based to virtual-machine-based implementations, for securing RT applications. It offers in addition an in-depth evaluation of these architectures, providing insights into how various TEE deployments influence the temporal compute and communication guarantees of RT systems. Peterson Yuhala, Christian Göttel, Jämes Ménétrey, Valerio Schiavoni, David Kozhaya, Pascal Felber |
ECRTS | 5 |
| 2024 | OneShot: View-Adapting Streamlined BFT Protocols with Trusted Execution EnvironmentsabstractByzantine fault-tolerance is arguably an expensive characteristic for protocols to support, especially when considering its overhead on message complexity, number of communication phases, and number of nodes for resilience. Various works in the literature have addressed optimizing one or more of these dimensions through the use of algorithmic optimizations, trusted execution environments, and streamlined view changes.The best achievable message complexity, resilience, and latency to date are respectively linear complexity, by ⌊(N − 1)/2⌋, and two communication phases attained Damysus (EuroSys’22), a streamlined hybrid protocol. This paper strictly advances the aforementioned state of the art results by introducing OneShot, a streamlined hybrid protocol that uses one communication phase in the normal case, and one or two phases otherwise. OneShot exploits the information nodes receive about the system to dynamically modify and adapt views. We prove that OneShot is safe and live, and moreover demonstrate through experimental evaluation that it improves throughput and latency by respectively up to 150% and 59% compared to the state of the art. Jeremie Decouchant, David Kozhaya, Vincent Rahli, Jiangshan Yu |
IPDPS | 2 |
| 2023 | Qualitative Analysis for Validating IEC 62443-4-2 Requirements in DevSecOpsabstractValidation of conformance to cybersecurity standards for industrial automation and control systems is an expensive and time consuming process which can delay the time to market. It is therefore crucial to introduce conformance validation stages into the continuous integration/continuous delivery pipeline of products. However, designing such conformance validation in an automated fashion is a highly non-trivial task that requires expert knowledge and depends upon available security tools, ease of integration into the DevOps pipeline, as well as support for IT and OT interfaces and protocols.This paper addresses the aforementioned problem focusing on the automated validation of ISA/IEC 62443-4-2 standard component requirements. We present an extensive qualitative analysis of the standard requirements and the current tooling landscape to perform validation. Our analysis demonstrates the coverage established by the currently available tools and sheds light on current gaps to achieve full automation and coverage. Furthermore, we showcase for every component requirement where in the CI/CD pipeline stage it is recommended to test it and the tools to do so. Christian Göttel, Maelle Kabir-Querrec, David Kozhaya, Thanikesavan Sivanthi, Ognjen Vukovic |
ETFA | 3 |
| 2022 | DAMYSUS: streamlined BFT consensus leveraging trusted componentsabstractRecently, streamlined Byzantine Fault Tolerant (BFT) consensus protocols, such as HotStuff, have been proposed as a means to circumvent the inefficient view-changes of traditional BFT protocols, such as PBFT. Several works have detailed trusted components, and BFT protocols that leverage them to tolerate a minority of faulty nodes and use a reduced number of communication rounds. Inspired by these works we identify two basic trusted services, respectively called the Checker and Accumulator services, which can be leveraged by streamlined protocols. Based on these services, we design Damysus, a streamlined protocol that improves upon HotStuff's resilience and uses less communication rounds. In addition, we show how the Checker and Accumulator services can be adapted to develop Chained-Damysus, a chained version of Damysus where operations are pipelined for efficiency. We prove the correctness of Damysus and Chained-Damysus, and evaluate their performance showcasing their superiority compared to previous protocols. Jeremie Decouchant, David Kozhaya, Vincent Rahli, Jiangshan Yu |
EuroSys | 2 |
| 2021 | Probabilistic and temporal failure detectors for solving distributed problems
Rachid Guerraoui, David Kozhaya, Yvonne-Anne Pignolet |
J. Parallel Distributed Comput. | 2 |
| 2021 | PISTIS: An Event-Triggered Real-Time Byzantine-Resilient Protocol SuiteabstractThe accelerated digitalisation of society along with technological evolution have extended the geographical span of cyber-physical systems. Two main threats have made the reliable and real-time control of these systems challenging: (i) uncertainty in the communication infrastructure induced by scale, and heterogeneity of the environment and devices; and (ii) targeted attacks maliciously worsening the impact of the above-mentioned communication uncertainties, disrupting the correctness of real-time applications. This article addresses those challenges by showing how to build distributed protocols that provide both real-time with practical performance, and scalability in the presence of network faults and attacks, in probabilistic synchronous environments. We provide a suite of real-time Byzantine protocols, which we prove correct, starting from a reliable broadcast protocol, called PISTIS, up to atomic broadcast and consensus. This suite simplifies the construction of powerful distributed and decentralized monitoring and control applications, including state-machine replication. Extensive empirical simulations showcase PISTIS's robustness, latency, and scalability. For example, PISTIS can withstand message loss (and delay) rates up to 50 percent in systems with 49 nodes and provides bounded delivery latencies in the order of a few milliseconds. David Kozhaya, Jeremie Decouchant, Vincent Rahli, Paulo Veríssimo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Byzantine Resilient Protocol for the IoTabstractWireless sensor networks (WSNs), often adhering to a single gateway architecture, constitute the communication backbone for many modern cyber-physical systems (CPSs). Consequently, fault-tolerance in CPS becomes a challenging task, especially when accounting for failures (potentially malicious) that incapacitate the gateway or disrupt the nodes-gateway communication, not to mention the energy, timeliness, and security constraints demanded by CPS domains. This paper aims at ameliorating the fault-tolerance of WSN-based CPS to increase system and data availability. To this end, we propose a replicated gateway architecture augmented with energy-efficient real-time Byzantine-resilient data communication protocols. At the sensors level, we introduce fault-tolerant trustful space-time protocol, a geographic routing protocol capable of delivering messages in an energy-efficient and timely manner to multiple gateways, even in the presence of voids caused by faulty and malicious sensor nodes. At the gateway level, we propose a multigateway synchronization protocol, which we call ByzCast, that delivers timely correct data to CPS applications, despite the failure or maliciousness of a number of gateways. We show, through extensive simulations, that our protocols provide better system robustness yielding an increased system and data availability while meeting CPS energy, timeliness, and security demands. Antônio Augusto Fröhlich, Roberto Milton Scheffel, David Kozhaya, Paulo Veríssimo |
IEEE Internet Things J. | 3 |
| 2019 | RT-ByzCast: Byzantine-Resilient Real-Time Reliable BroadcastabstractToday's cyber-physical systems face various impediments to achieving their intended goals, namely, communication uncertainties and faults, relative to the increased integration of networked and wireless devices, hinder the synchronism needed to meet real-time deadlines. Moreover, being critical, these systems are also exposed to significant security threats. This threat combination increases the risk of physical damage. This paper addresses these problems by studying how to build the first real-time Byzantine reliable broadcast protocol (RTBRB) tolerating network uncertainties, faults, and attacks. Previous literature describes either real-time reliable broadcast protocols, or asynchronous (non real-time) Byzantine ones. We first prove that it is impossible to implement RTBRB using traditional distributed computing paradigms, e.g., where the error/failure detection mechanisms of processes are decoupled from the broadcast algorithm itself, even with the help of the most powerful failure detectors. We circumvent this impossibility by proposing RT-ByzCast, an algorithm based on aggregating digital signatures in a sliding time-window and on empowering processes with self-crashing capabilities to mask and bound losses. We show that RT-ByzCast (i) operates in real-time by proving that messages broadcast by correct processes are delivered within a known bounded delay, and (ii) is reliable by demonstrating that correct processes using our algorithm crash themselves with a negligible probability, even with message loss rates as high as 60 percent. David Kozhaya, Jeremie Decouchant, Paulo Veríssimo |
IEEE Trans. Computers | 1 |
| 2019 | RepuCoin: Your Reputation Is Your PowerabstractExisting proof-of-work cryptocurrencies cannot tolerate attackers controlling more than 50 percent of the network's computing power at any time, but assume that such a condition happening is “unlikely”. However, recent attack sophistication, e.g., where attackers can rent mining capacity to obtain a majority of computing power temporarily, render this assumption unrealistic. This paper proposes RepuCoin, the first system to provide guarantees even when more than 50 percent of the system's computing power is temporarily dominated by an attacker. RepuCoin physically limits the rate of voting power growth of the entire system. In particular, RepuCoin defines a miner's power by its `reputation', as a function of its work integrated over the time of the entire blockchain, rather than through instantaneous computing power, which can be obtained relatively quickly and/or temporarily. As an example, after a single year of operation, RepuCoin can tolerate attacks compromising 51 percent of the network's computing resources, even if such power stays maliciously seized for almost a whole year. Moreover, RepuCoin provides better resilience to known attacks, compared to existing proof-of-work systems, while achieving a high throughput of 10000 transactions per second (TPS). Jiangshan Yu, David Kozhaya, Jeremie Decouchant, Paulo Veríssimo |
IEEE Trans. Computers | 2 |
| 2018 | You Only Live Multiple Times: A Blackbox Solution for Reusing Crash-Stop Algorithms In Realistic Crash-Recovery SettingsabstractDistributed agreement-based algorithms are often specified in a crash-stop asynchronous model augmented by Chandra and Toueg's unreliable failure detectors. In such models, correct nodes stay up forever, incorrect nodes eventually crash and remain down forever, and failure detectors behave correctly forever eventually, However, in reality, nodes as well as communication links both crash and recover without deterministic guarantees to remain in some state forever. In this paper, we capture this realistic temporary and probabilitic behaviour in a simple new system model. Moreover, we identify a large algorithm class for which we devis a property-preserving transformation. Using this transformation, many algorithms written for the asynchronous crash-stop model run correctly and unchanged in real systems. David Kozhaya, Ognjen Maric, Yvonne-Anne Pignolet |
OPODIS | 1 |
| 2016 | Never Say Never - Probabilistic and Temporal Failure DetectorsabstractThe failure detector approach for solving distributed computing problems has been celebrated for its modularity. This approach allows the construction of algorithms using abstract failure detection mechanisms, defined by axiomatic properties, as building blocks. The minimal synchrony assumptions on communication, which enable to implement the failure detection mechanism, are studied separately. Such synchrony assumptions are typically expressed as eventual guarantees that need to hold, after some point in time, forever and deterministically. But in practice, they never do. Synchrony assumptions may hold only probabilistically and temporarily. In this paper, we study failure detectors in a realistic distributed system N, with asynchrony inflicted by probabilistic synchronous communication. We address the following paradox: an implementation of "consensus with probability 1" is possible in N without using randomness in the algorithm itself, while an implementation of "◇S with probability 1" is impossible to achieve in N (◇S being the weakest failure detector to solve the consensus problem and many equivalent problems). We circumvent this paradox by introducing a new failure detector ◇S*, a variant of ◇S with probabilistic and temporal accuracy. We prove that ◇S* is implementable in N and we provide an optimal ◇S* algorithm. Interestingly, we show that ◇S* can replace ◇S, in several existing deterministic consensus algorithms using ◇S, to yield an algorithm that solves "consensus with probability 1". In fact, we show that such result holds for all decisive problems (not only consensus) and also for failure detector ◇P (not only ◇S). The resulting algorithms combine the modularity of distributed computing practices with the practicality of networking ones. Dacfey Dzung, Rachid Guerraoui, David Kozhaya, Yvonne-Anne Pignolet |
IPDPS | 3 |
| 2016 | Right on Time Distributed Shared MemoryabstractThe demand for real-time data storage in distributed control systems (DCSs) is growing. Yet, providing real-time DCS guarantees is challenging, especially when more and more sensor and actuator devices are connected to industrial plants and message loss needs to be taken into account.In this paper, we investigate how to build a shared memory abstraction for DCSs as a first step towards implementing different shared storage systems in a DCS context. We first prove that, in the presence of host crashes and message losses, the necessary guarantees of such an abstraction are impossible to implement using a traditional approach that has no access to the internals of existing DCS services, e.g., a modular approach where algorithms are built on top of existing software blocks like failure detectors. We propose a white-box approach that utilizes messages of existing services in any DCS as the sole means of communication. More precisely, we present TapeWorm, an algorithm that attaches itself to the heartbeat messages of the failure detector component in DCSs. We prove that TapeWorm implements the desired shared memory guarantees for applications running on a DCS. We also analyze the performance of TapeWorm and we showcase ways of adapting TapeWorm to various application needs and workloads. Rachid Guerraoui, David Kozhaya, Yvonne-Anne Pignolet |
RTSS | 2 |
| 2016 | Who's On Board?: Probabilistic Membership for Real-Time Distributed Control SystemsabstractTo increase their dependability, distributed control systems (DCSs) need to agree in real time about which hosts have crashed, i.e., they need a real-time membership service. In this paper, we prove that such a service cannot be implemented deterministically if, besides host crashes, communication can also fail. We define implementable probabilistic variants of membership properties, which constitute what we call a synchronous membership service (SYMS). We present an algorithm, ViewSnoop, that implements SYMS with high-probability. We implement, deploy and evaluate ViewSnoop analytically as well as experimentally, within an industrial DCS framework. We show that ViewSnoop significantly improves the dependability of DCSs compared to membership schemes based on classic heartbeats, at low additional cost. Moreover, ViewSnoop distinguishes, with high probability, host crashes from message losses, enabling DCSs to counteract losses better than existing approaches. Rachid Guerraoui, David Kozhaya, Manuel Oriol, Yvonne-Anne Pignolet |
SRDS | 2 |
| 2015 | To Transmit Now or Not to Transmit NowabstractGiven an unreliable communication link, this paper studies how to build, in an energy-efficient manner, a reliable communication service that is synchronous with high probability. We consider a Partially Observable Markov Decision Process (POMDP) setting in which a communication link's transmission quality: (i) changes according to a classic Markovian model and (ii) can be only partially observed, through feedback relative to previous transmissions. We perform a thorough analysis under several variations of Ack/Nack feedback mechanisms. Despite the general intractability of POMDPs, we prove that our communication service, under reliable feedback, can be inexpensively implemented. We obtain closed form solutions specifying when to transmit over the link, which allows to derive an energy-optimal implementation. We also analyse the impact of lossy feedback on implementing our communication service. Considering multiple lossy feedback mechanisms, we show that an easily implementable structure for our communication service can also be obtained, depending on the feedback mechanism itself. Dacfey Dzung, Rachid Guerraoui, David Kozhaya, Yvonne-Anne Pignolet |
SRDS | 3 |