Pavlos Aimoniotis

dblp:287/4417 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-6602-1988ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RXT: RefleXive Address Translation for Pointer-Chasing Workloads
abstract
With increasingly irregular memory access patterns in Indirect Memory Access (IMA) and Graph Processing (GP) applications, Virtual Address Translation (VAT) not only has become a major performance bottleneck, but also a key contributor to high power consumption in out-of-order (OoO) multi-core processor pipeline. We find that this problem can be attributed to a large extent to an inefficient utilization of TLBs in pointer-chasing instructions (i.e. load-to-load), which constitute a significant portion of workloads in IMA/GP. Conventionally, VATs for pointer-chasing are performed via a series of TLB look-ups: one for each load in the pointer chain. However, this approach is sub-optimal as it overlooks the correlations between pointers, missing opportunities to perform VATs through more costeffective calculations instead of relying on the more expensive TLB look-ups. This work draws on the insight that in pointer-chasing workloads, the physical page holding an upstream pointer is often at a fixed distance to the physical page where a downstream pointer resides. Building on this insight, we introduce RefleXive address Translation (RXT), which encodes the physical page distance (termed PageDist) between an upstream and downstream pointer into the unused upper 16 bits of the upstream pointer's virtual address. Consequently, RXT can compute the translation of a pointer by directly adding the PageDist to its residing address (the physical address where the pointer is stored). This process can be recursively applied throughout the pointer-chasing sequence, transforming address translation from a sequence of TLB accesses to computes. Through gem5 full-system simulation, RXT reduces TLB energy consumption for pointer-chasing applications by an average of 41.48 % (up to 62.50 %) and decreases overall core power density by 4.65 % (up to 13.09 %), without compromising application execution time. In fact, RXT even improves program runtime by an average of 0.71 % (up to 2.09 %).
Rashid Aligholipour, Pavlos Aimoniotis, Stefanos Kaxiras, Yuan Yao 0009
IPDPS2
2024 JANUS: A Simple and Efficient Speculative Defense using Reinforcement Learning
abstract
Speculative execution and the emergence of Spectre attacks have forced architects to rethink how microprocessors are designed. Several approaches aim to close this security vulnerability while trying to minimize performance degradation, often involving complex and sophisticated mechanisms. These strategies typically entail substantial modifications to the processor core and memory hierarchy, which ultimately inhibit their adoption in real designs.In this work, we leverage two of the simplest speculative defense ideas, NDA and DoM, that can co-exist in the same core, and we apply a simple form of Reinforcement Learning (RL) to select the most effective mechanism, as the underlying processor defense, for a window of execution. NDA forbids the propagation of a potential secret to subsequent instructions while DoM prohibits the creation of observable timing differences in the cache. We observe that their impact on different applications can vary significantly, but, often, they can complement each other within the same application. However, our investigation also reveals vulnerabilities in previous proposals that try to combine these secure speculation schemes into one. We demonstrate an attack scenario that violates the security of the combined scheme and we present the conditions that must hold to safely combine them. Lastly, while the cost and complexity of reinforcement learning may seem inordinately high for microarchitectural implementations, we build on recent research that demonstrates remarkably lightweight solutions, provided that the action space is small.We present JANUS, a lightweight architecture leveraging an RL agent based on a two-armed bandit algorithm. JANUS selects the optimal, performance-wise, defense mechanism that protects the processor within a specific time window. We evaluate JANUS on SPEC2017 benchmark suite and find that it outperforms NDA by +4.9%, STT (a more sophisticated and complex scheme that uses taint tracking) by +1%, and DoM by +2.6%. Further, when a state-of-the-art address-prediction optimization (Doppelganger Loads) is employed on top of the baseline defenses, NDA and DoM, JANUS still outperforms the former by +2.3%, and the latter by +0.3%. When evaluated with the older SPEC2006 benchmark suite, JANUS outperforms all schemes by +4.7% on average, with a maximum of +8.2% over DoM. JANUS achieves these results with a meager storage overhead of just 16 bytes and a complexity-effective design.
Pavlos Aimoniotis, Stefanos Kaxiras
SBAC-PAD1
2023 Doppelganger Loads: A Safe, Complexity-Effective Optimization for Secure Speculation Schemes
abstract
Speculative side-channel attacks have forced computer architects to rethink speculative execution. Effectively preventing microarchitectural state from leaking sensitive information will be a key requirement in future processor design.
Amund Bergland Kvalsvik, Pavlos Aimoniotis, Stefanos Kaxiras, Magnus Själander
ISCA2
2023 ReCon: Efficient Detection, Management, and Use of Non-Speculative Information Leakage
abstract
In a speculative side-channel attack, a secret is improperly accessed and then leaked by passing it to a transmitter instruction. Several proposed defenses effectively close this security hole by either delaying the secret from being loaded or propagated, or by delaying dependent transmitters (e.g., loads) from executing when fed with tainted input derived from an earlier speculative load. This results in a loss of memory-level parallelism and performance.
Pavlos Aimoniotis, Amund Bergland Kvalsvik, Magnus Själander, Stefanos Kaxiras
MICRO1