Christos Sakalis

dblp:153/5786 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
2since 2021 · last 2023
0000-0003-4172-8607ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
3 papers
Hardware security and side channels · 74% Authentication and access control · 26%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 82% Memory systems · 18%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
speculative execution
0.822020
Understanding Selective Delay as a Method for Efficient Secure Speculative Execution · IEEE Trans. Computers 2020
Efficient invisible speculative execution through selective delay and value prediction · ISCA 2019
Hardware security and side channels › microarchitectural attacks
microarchitectural replay attack
0.712023
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks · ACM Trans. Archit. Code Optim. 2023
Authentication and access control
replay attack prevention
0.712023
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks · ACM Trans. Archit. Code Optim. 2023
Processor architecture and microarchitecture › speculative execution
speculative execution security
0.712023
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks · ACM Trans. Archit. Code Optim. 2023
Hardware security and side channels › microarchitectural attacks › transient execution attack
speculative execution attack
0.412020
Understanding Selective Delay as a Method for Efficient Secure Speculative Execution · IEEE Trans. Computers 2020
Hardware security and side channels › secure speculation
speculative side-channel defense
0.412020
Understanding Selective Delay as a Method for Efficient Secure Speculative Execution · IEEE Trans. Computers 2020
Processor architecture and microarchitecture › speculative execution
secure speculation
0.412020
Understanding Selective Delay as a Method for Efficient Secure Speculative Execution · IEEE Trans. Computers 2020
Memory systems
cache coherence
0.312017
Efficient Self-Invalidation/Self-Downgrade for Critical Sections with Relaxed Semantics · IEEE Trans. Parallel Distributed Syst. 2017
Hardware security and side channels
side-channel attack
0.212023
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks · ACM Trans. Archit. Code Optim. 2023
Memory systems › memory access optimization
memory-level parallelism
0.112020
Understanding Selective Delay as a Method for Efficient Secure Speculative Execution · IEEE Trans. Computers 2020
Hardware security and side channels › microarchitectural attacks › transient execution attack › speculative execution attack
spectre and meltdown
0.112019
Efficient invisible speculative execution through selective delay and value prediction · ISCA 2019
Concurrent programming › synchronization
critical sections
0.112017
Efficient Self-Invalidation/Self-Downgrade for Critical Sections with Relaxed Semantics · IEEE Trans. Parallel Distributed Syst. 2017
Concurrent programming
synchronization
0.112017
Efficient Self-Invalidation/Self-Downgrade for Critical Sections with Relaxed Semantics · IEEE Trans. Parallel Distributed Syst. 2017

Methods — techniques the papers use, named apart from their topics

value prediction · 1.6instruction squash tracking · 1.3delay-on-miss · 0.9selective delay · 0.8relaxed atomic operations · 0.6directory protocol · 0.6
YearPublicationVenuePosition
2023 Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks
abstract
MicroScope and other similar microarchitectural replay attacks take advantage of the characteristics of speculative execution to trap the execution of the victim application in a loop, enabling the attacker to amplify a side-channel attack by executing it indefinitely. Due to the nature of the replay, it can be used to effectively attack software that are shielded against replay, even under conditions where a side-channel attack would not be possible (e.g., in secure enclaves). At the same time, unlike speculative side-channel attacks, microarchitectural replay attacks can be used to amplify the correct path of execution, rendering many existing speculative side-channel defenses ineffective. In this work, we generalize microarchitectural replay attacks beyond MicroScope and present an efficient defense against them. We make the observation that such attacks rely on repeated squashes of so-called “replay handles” and that the instructions causing the side-channel must reside in the same reorder buffer window as the handles. We propose Delay-on-Squash, a hardware-only technique for tracking squashed instructions and preventing them from being replayed by speculative replay handles. Our evaluation shows that it is possible to achieve full security against microarchitectural replay attacks with very modest hardware requirements while still maintaining 97% of the insecure baseline performance.
Christos Sakalis, Stefanos Kaxiras, Magnus Själander
ACM Trans. Archit. Code Optim.1
2021 Splash-4: Improving Scalability with Lock-Free Constructs
abstract
Over the past three decades, the parallel applications of the Splash-2 benchmark suite have been instrumental in advancing multiprocessor research. Recently, the Splash-3 benchmarks eliminated performance bugs, data races, and improper synchronization that plagued Splash-2 benchmarks after the definition of the C memory model. In this work, we revisit the Splash-3 benchmarks and adapt them for contemporary architectures with atomic operations and lock-free constructs. With our changes, we improve the scalability of most benchmarks for up to 32 and 64 cores, showing an improvement of up to 9x in actual machines, and up to 5x in simulation, over the unmodified Splash-3 benchmarks. To denote the substantive nature of the improvements in the Splash-3 benchmarks and to re-introduce them in contemporary research, we refer to the new collection as Splash-4.
Eduardo José Gómez-Hernández, Ruixiang Shao, Christos Sakalis, Stefanos Kaxiras, Alberto Ros 0001
ISPASS3
2020 Clearing the Shadows: Recovering Lost Performance for Invisible Speculative Execution through HW/SW Co-Design
abstract
Out-of-order processors heavily rely on speculation to achieve high performance, allowing instructions to bypass other slower instructions in order to fully utilize the processor's resources. Speculatively executed instructions do not affect the correctness of the application, as they never change the architectural state, but they do affect the micro-architectural behavior of the system. Until recently, these changes were considered to be safe but with the discovery of new security attacks that misuse speculative execution to leak secrete information through observable micro-architectural changes (so called side-channels), this is no longer the case. To solve this issue, a wave of software and hardware mitigations have been proposed, the majority of which delay and/or hide speculative execution until it is deemed to be safe, trading performance for security. These newly enforced restrictions change how speculation is applied and where the performance bottlenecks appear, forcing us to rethink how we design and optimize both the hardware and the software.
Kim-Anh Tran, Christos Sakalis, Magnus Själander, Alberto Ros 0001, Stefanos Kaxiras, Alexandra Jimborean
PACT2
2020 Evaluating the Potential Applications of Quaternary Logic for Approximate Computing
abstract
There exist extensive ongoing research efforts on emerging atomic-scale technologies that have the potential to become an alternative to today’s complementary metal--oxide--semiconductor technologies. A common feature among the investigated technologies is that of multi-level devices, particularly the possibility of implementing quaternary logic gates and memory cells. However, for such multi-level devices to be used reliably, an increase in energy dissipation and operation time is required. Building on the principle of approximate computing, we present a set of combinational logic circuits and memory based on multi-level logic gates in which we can trade reliability against energy efficiency. Keeping the energy and timing constraints constant, important data are encoded in a more robust binary format while error-tolerant data are encoded in a quaternary format. We analyze the behavior of the logic circuits when exposed to transient errors caused as a side effect of this encoding. We also evaluate the potential benefit of the logic circuits and memory by embedding them in a conventional computer system on which we execute jpeg, sobel, and blackscholes approximately. We demonstrate that blackscholes is not suitable for such a system and explain why. However, we also achieve dynamic energy reductions of 10% and 13% for jpeg and sobel, respectively, and improve execution time by 38% for sobel, while maintaining adequate output quality.
Christos Sakalis, Alexandra Jimborean, Stefanos Kaxiras, Magnus Själander
ACM J. Emerg. Technol. Comput. Syst.1
2020 Understanding Selective Delay as a Method for Efficient Secure Speculative Execution
abstract
Since the introduction of Meltdown and Spectre, the research community has been tirelessly working on speculative side-channel attacks and on how to shield computer systems from them. To ensure that a system is protected not only from all the currently known attacks but also from future, yet to be discovered, attacks, the solutions developed need to be general in nature, covering a wide array of system components, while at the same time keeping the performance, energy, area, and implementation complexity costs at a minimum. One such solution is our own delay-on-miss, which efficiently protects the memory hierarchy by i) selectively delaying speculative load instructions and ii) utilizing value prediction as an invisible form of speculation. In this article we dive deeper into delay-on-miss, offering insights into why and how it affects the performance of the system. We also reevaluate value prediction as an invisible form of speculation. Specifically, we focus on the implications that delaying memory loads has in the memory level parallelism of the system and how this affects the value predictor and the overall performance of the system. We present new, updated results but more importantly, we also offer deeper insight into why delay-on-miss works so well and what this means for the future of secure speculative execution.
Christos Sakalis, Stefanos Kaxiras, Alberto Ros 0001, Alexandra Jimborean, Magnus Själander
IEEE Trans. Computers1
2019 Ghost loads: what is the cost of invisible speculation?
abstract
Speculative execution is necessary for achieving high performance on modern general-purpose CPUs but, starting with Spectre and Meltdown, it has also been proven to cause severe security flaws. In case of a misspeculation, the architectural state is restored to assure functional correctness but a multitude of microarchitectural changes (e.g., cache updates), caused by the speculatively executed instructions, are commonly left in the system. These changes can be used to leak sensitive information, which has led to a frantic search for solutions that can eliminate such security flaws. The contribution of this work is an evaluation of the cost of hiding speculative side-effects in the cache hierarchy, making them visible only after the speculation has been resolved. For this, we compare (for the first time) two broad approaches: i) waiting for loads to become non-speculative before issuing them to the memory system, and ii) eliminating the side-effects of speculation, a solution consisting of invisible loads (Ghost loads) and performance optimizations (Ghost Buffer and Materialization). While previous work, InvisiSpec, has proposed a similar solution to our latter approach, it has done so with only a minimal evaluation and at a significant performance cost. The detailed evaluation of our solutions shows that: i) waiting for loads to become non-speculative is no more costly than the previously proposed InvisiSpec solution, albeit much simpler, non-invasive in the memory system, and stronger security-wise; ii) hiding speculation with Ghost loads (in the context of a relaxed memory model) can be achieved at the cost of 12% performance degradation and 9% energy increase, which is significantly better that the previous state-of-the-art solution.
Christos Sakalis, Mehdi Alipour, Alberto Ros 0001, Alexandra Jimborean, Stefanos Kaxiras, Magnus Själander
CF1
2019 Efficient invisible speculative execution through selective delay and value prediction
abstract
Speculative execution, the base on which modern high-performance general-purpose CPUs are built on, has recently been shown to enable a slew of security attacks. All these attacks are centered around a common set of behaviors: During speculative execution, the architectural state of the system is kept unmodified, until the speculation can be verified. In the event that a misspeculation occurs, then anything that can affect the architectural state is reverted (squashed) and re-executed correctly. However, the same is not true for the microarchitectural state. Normally invisible to the user, changes to the microarchitectural state can be observed through various side-channels, with timing differences caused by the memory hierarchy being one of the most common and easy to exploit. The speculative side-channels can then be exploited to perform attacks that can bypass software and hardware checks in order to leak information. These attacks, out of which the most infamous are perhaps Spectre and Meltdown, have led to a frantic search for solutions.
Christos Sakalis, Stefanos Kaxiras, Alberto Ros 0001, Alexandra Jimborean, Magnus Själander
ISCA1
2017 Efficient Self-Invalidation/Self-Downgrade for Critical Sections with Relaxed Semantics
abstract
Cache coherence protocols based on self-invalidation allow simpler hardware implementation compared to traditional write-invalidation protocols, by relying on data-race-free semantics and applying self-invalidation on synchronization points. Their simplicity lies in the absence of invalidation traffic. This eliminates the need to track readers in a directory, and reduces the number of transient protocol states. Similarly, the use of self-downgrade on synchronization eliminates directory indirection, and hence the need to track writers in a directory. These protocols, effectively without a directory, have the potential to reduce area, energy consumption, and complexity, without sacrificing performance-provided, that self-invalidation and self-downgrade are performed prudently. In this work we examine how self-invalidation and self-downgrade are performed in relation to atomicity and ordering. We show that self-invalidation and self-downgrade do not need to be applied conservatively, as so far implemented. Our key observation is that, often, critical sections which are not ordered in time, are intended to provide only atomicity and not thread synchronization. We thus propose a new type of self-invalidation, forward-self-invalidation (FSI), which invalidates solely data that are going to be accessed inside a critical section. Based on the same reasoning, we propose a new type of self-downgrade, forward self-downgrade (FSD), also restricted to writes in critical sections. Finally, we define the semantics of locks using FSI and FSD, which resemble the semantics of relaxed atomic operations in C++. Our evaluation for 64-core multiprocessors shows significant improvements using the proposed FSI and FSD-where applicable-in Splash-3 and PARSEC benchmarks, over a directory-based protocol (17.1 percent in execution time and 33.9 percent in energy consumption) and also over a state-of-the-art self-invalidation/self-downgrade protocol (7.6 percent in execution time and 9.1 percent in energy consumption), while still retaining the design simplicity of the protocol.
Alberto Ros 0001, Carl Leonardsson, Christos Sakalis, Stefanos Kaxiras
IEEE Trans. Parallel Distributed Syst.3
2016 POSTER: Efficient Self-Invalidation/Self-Downgrade for Critical Sections with Relaxed Semantics
abstract
Cache coherence protocols based on self-invalidation allow simpler hardware implementation compared to traditional write-invalidation protocols, by relying on data-race-free semantics and applying self-invalidation and self-downgrade on synchronization points. This work examines how self-invalidation and self-downgrade are performed in relation to atomicity and ordering and shows that they do not need to be applied conservatively, as so far implemented. Our key observation is that, often, critical sections which are not ordered in time, are intended to provide only atomicity but not thread synchronization.
Alberto Ros 0001, Carl Leonardsson, Christos Sakalis, Stefanos Kaxiras
PACT3
2016 Splash-3: A properly synchronized benchmark suite for contemporary research
abstract
Benchmarks are indispensable in evaluating the performance implications of new research ideas. However, their usefulness is compromised if they do not work correctly on a system under evaluation or, in general, if they cannot be used consistently to compare different systems. A well-known benchmark suite of parallel applications is the Splash-2 suite. Since its creation in the context of the DASH project, Splash-2 benchmarks have been widely used in research. However, Splash-2 was released over two decades ago and does not adhere to the recent C memory consistency model. This leads to unexpected and often incorrect behavior when some Splash-2 benchmarks are used in conjunction with contemporary compilers and hardware (simulated or real). Most importantly, we discovered critical performance bugs that may question some of the reported benchmark results. In this work, we analyze the Splash-2 benchmarks and expose data races and related performance bugs. We rectify the problematic benchmarks and evaluate the resulting performance. Our work contributes to the community a new sanitized version of the Splash-2 benchmarks, called the Splash-3 benchmark suite.
Christos Sakalis, Carl Leonardsson, Stefanos Kaxiras, Alberto Ros 0001
ISPASS1