Ali Hajiabadi

dblp:276/6511 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-3219-7544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MIRZA: Efficiently Mitigating Rowhammer with Randomization and ALERT
abstract
In-DRAM Rowhammer mitigation requires three resources: space (to track aggressor rows), time (to perform mitigation), and energy (to refresh victim rows). An ideal in-DRAM mitigation must minimize all three overheads. Recent randomized trackers, such as MINT, can perform tracking with negligible storage overheads. However, they perform mitigation proactively and frequently, which incurs significant performance and energy overheads at low thresholds. Recently, JEDEC introduced Per-Row Activation Counters (PRAC) and ALERT Back Off (ABO) protocol to obtain the time for mitigation reactively, as needed. While PRAC+ABO minimizes the time and energy overheads of mitigation, PRAC incurs significant changes to the DRAM array and significant performance overhead (6.5% on average) due to increased memory timings to update the PRAC counters. Our goal is to develop an efficient in-DRAM mitigation that has low storage, performance, and energy overheads. Our paper proposes MIRZA, the first low-cost reactive in-DRAM mitigation. MIRZA relies on MINT to track aggressor rows. However, instead of proactively doing mitigation at regular intervals (via REF or RFM), MIRZA uses ABO to reactively obtain the time required for mitigation. To avoid frequent ABO, MIRZA employs Coarse-Grained Filtering to disable mitigations if the activation count is below a certain Filtering Threshold. To tolerate a threshold of 1K, MIRZA requires a storage overhead of only 196 bytes of SRAM per bank. Compared to MINT, MIRZA reduces the mitigation overheads by 28.5×. Compared to PRAC, MIRZA has 45× lower area overheads and negligible slowdown (0.36% average slowdown vs. 6.5% for PRAC).
Hritvik Taneja, Ali Hajiabadi, Michele Marazzi, Kaveh Razavi, Moinuddin K. Qureshi
HPCA2
2026 VMSCAPE: Exposing and Exploiting Incomplete Branch Predictor Isolation in Cloud Environments
Jean-Claude Graf, Sandro Rüegge, Ali Hajiabadi, Kaveh Razavi
SP3
2025 CHaRM: Checkpointed and Hashed Counters for Flexible and Efficient Rowhammer Mitigation
abstract
Despite efforts by DRAM vendors to mitigate Rowhammer, it is still a potent attack vector. CPU vendors are reluctant to deploy deterministic mitigations against Rowhammer due to the high cost that needs to be paid for the most vulnerable DRAM device, even though an average DRAM device is considerably less vulnerable. The main reason for this high cost is the need to track an increasing number of aggressor rows with the worsening Rowhammer threshold. Our proposed in-CPU mitigation, called CHaRM, breaks this dependency by efficiently mapping a large number of rows to a fixed number of hashed counters. Since multiple rows are now mapped to a limited number of counters, collisions can occur. To avoid excessive mitigative refreshes upon collisions, CHaRM deploys a checkpointing mechanism that saves the state of rows evicted from the table. When a row is activated again, CHaRM restores its checkpointed value and resumes tracking. Our evaluation shows that CHaRM incurs negligible slowdown, below 1% across all Rowhammer thresholds, while improving area, power, and energy by 3.8x, 4.4x, and 8.2x, respectively, for Rowhammer threshold of 1K compared to the state of the art.
Ali Hajiabadi, Michele Marazzi, Kaveh Razavi
CCS1
2025 Cassandra: Efficient Enforcement of Sequential Execution for Cryptographic Programs
abstract
Constant-time programming is a widely deployed approach to harden cryptographic programs against side channel attacks.However, modern processors often violate the underlying assumptions of standard constant-time policies by transiently executing unintended paths of the program.Despite many solutions proposed, addressing control flow misspeculations in an efficient way without losing performance is an open problem.In this work, we propose Cassandra, a novel hardware/software mechanism to enforce sequential execution for constant-time cryptographic code in a highly efficient manner.Cassandra explores the radical design point of disabling the branch predictor and recording-and-replaying sequential control flow of the program.Two key insights that enable our design are that (1) the sequential control flow of a constant-time program is mostly static over different runs, and (2) cryptographic programs are loop-intensive and their control flow patterns repeat in a highly compressible way.These insights allow us to perform an upfront branch analysis that significantly compresses control flow traces.We add a small component to a typical processor design, the Branch Trace Unit, to store compressed traces and determine fetch redirections according to the sequential model of the program.Despite providing a strong security guarantee, Cassandra counterintuitively provides an average 1.85% speedup compared to an unsafe baseline processor, mainly due to enforcing near-perfect fetch redirections.
Ali Hajiabadi, Trevor E. Carlson
ISCA1
2025 One Flew over the Stack Engine's Nest: Practical Microarchitectural Attacks on the Stack Engine
abstract
Security research on modern CPUs has raised numerous concerns in recent years.These security issues stem from classic microarchitectural optimizations designed decades ago, without consideration for security.Stack pointer tracking, also known as the stack engine in recent CPUs, is one such optimization.To investigate the security implications of the stack engine, we reverse engineer its operational details on a number of recent Intel and AMD CPUs for the first time.Our results show the particular microarchitecturedependent behaviors of the stack engine, such as the conditions under which it needs to synchronize the stack pointer values with the backend.Using these results, we build three primitives called Direct Underflow, Sync+Reload and Prime+Sync+Probe that enable information leakage through the stack engine under different conditions.We use these primitives in the construction of various covert and side-channel attacks, leaking sensitive patient records from a widely-used JSON library as an example.Our mitigation efforts reveal that recent AMD Zen 4 and Zen 5 CPUs include undocumented chicken bits which allow enabling or disabling the stack engine.Using these bits to disable the stack engine, we measure 3.98% and 3.94% slowdown using SPEC CPU2017 on Zen 4 and Zen 5, respectively, prompting the need to consider more secure designs for the stack engine in future CPUs which we also discuss. CCS Concepts• Security and privacy → Side-channel analysis and countermeasures.
Silvan Niederer, Sandro Rüegge, Ali Hajiabadi, Kaveh Razavi
MICRO3
2025 PARADISE: Criticality-Aware Instruction Reordering for Power Attack Resistance
abstract
Power side-channel attacks exploit the correlation of power consumption with the instructions and data being processed to extract secrets from a device (e.g., cryptographic keys). Prior work primarily focused on protecting small embedded micro-controllers and in-order processors rather than high-performance, out-of-order desktop and server CPUs. In this article, we present Paradise , a general-purpose out-of-order processor with always-on protection, that implements a novel dynamic instruction scheduler to provide obfuscated execution and mitigate power analysis attacks. To achieve this, we exploit the time between operand availability of critical instructions ( slack ) and create high-performance random schedules. Further, we highlight the dangers of using incorrect adversarial assumptions, which can often lead to a false sense of security. Therefore, we perform an extended security analysis on AES-128 using different levels of adversaries, from basic to advanced, including a convolution neural networks–based attack. Our advanced security evaluation assumes a strong adversary with full knowledge of the countermeasure and demonstrates a significant security improvement of 556 × when combined with Boolean Masking over a baseline only protected by masking and 62,500× over an unprotected baseline. The resulting overhead in performance, power, and area of Paradise is 3.2%, 1.2%, and 0.8% respectively. 1
Yun Chen 0004, Ali Hajiabadi, Romain Poussier, Yaswanth Tavva, Andreas Diavastos, Shivam Bhasin, Trevor E. Carlson
ACM Trans. Archit. Code Optim.2
2024 Levioso: Efficient Compiler-Informed Secure Speculation
abstract
Spectre-type attacks have exposed a major class of vulnerabilities arising from speculative execution of instructions, the main performance enabler of modern CPUs. These attacks speculatively leak secrets that have been either speculatively loaded (seen in sand-boxed programs) or non-speculatively loaded (seen in constant-time programs). Various hardware-only defenses have been proposed to mitigate both speculative and non-speculative secrets via all potential transmission channels. However, limited program knowledge is exposed to the hardware and these solutions conservatively restrict the execution of all instructions that can potentially leak.
Ali Hajiabadi, Archit Agarwal, Andreas Diavastos, Trevor E. Carlson
DAC1
2024 Conjuring: Leaking Control Flow via Speculative Fetch Attacks
abstract
In this work, we propose a new attack called Conjuring that exploits one of the main features of CPUs' frontend: speculative fetch of instructions. We show that the Pattern History Table (PHT) in modern CPUs are a great channel to learn and leak control flow of victim applications. Unlike prior work, Conjuring does not require that one primes the PHT or interferes with the victim execution enabling a realistic and unprivileged attacker to leak control flow information. By improving the branch predictors, our attack becomes even more serious and practical. We demonstrate the feasibility of our attack on different existing Intel, AMD, and Apple CPUs.
Ali Hajiabadi, Trevor E. Carlson
DAC1
2024 GADGETSPINNER: A New Transient Execution Primitive Using the Loop Stream Detector
abstract
Transient execution attacks constitute a major class of attacks affecting all modern out-of-order CPUs. These attacks exploit transient execution windows (i.e., the instructions that execute but never commit) to leak confidential information from victims. Existing attacks either rely on branch mispredictions, incorrect memory speculation, or deferred exception handling to create transient windows. In this work, we introduce a new transient execution primitive, called GADGETSPINNER. We exploit the Loop Stream Detector (LSD) in Intel processors to perform out-of-loop-bounds execution and perform illegal operations. Our key observation is that the LSD holds on to an old copy of branch predictions from the first iteration of the loop and keeps using this copy until a branch misprediction occurs, i.e., advances beyond the loop bound. We exploit the delay between the speculative iteration of the loop and when the branch misprediction is resolved. In this paper, we analyze the transient execution of the LSD and perform end-to-end attacks to (1) perform illegal reads from protected memory regions, (2) bypass Intel SGX and extract the weights of a trained CNN model in DNNL library, (3) break Kernel ASLR (KASLR), and finally (4) perform cross-core/cross-process attacks. We also show that many defenses for prior transient execution attacks, like secure Branch Prediction Unit (BPU) designs, fail to protect against GADGETSPINNER.
Yun Chen 0004, Ali Hajiabadi, Trevor E. Carlson
HPCA2
2024 PREFETCHX: Cross-Core Cache-Agnostic Prefetcher-based Side-Channel Attacks
abstract
In this paper, we reveal the existence of a new class of prefetcher, the XPT prefetcher, in modern Intel processors which has never been officially detailed. It speculatively issues a load, bypassing last-level cache (LLC) lookups, when it predicts that a load request will result in an LLC miss. We demonstrate that XPT prefetcher is shared among different cores, which enables an attacker to build cross-core side-channel and covertchannel attacks. We propose PREFETCHX, a cross-core attack mechanism, to leak users’ sensitive data and activities. We empirically demonstrate that PREFETCHX can be used to extract private keys of real-world RSA applications. Furthermore, we show that PREFETCHX can enable side-channel attacks that can monitor keystrokes and network traffic patterns of users. Our two cross-core covert-channel attacks also see a low error rate and a 122KiB/s maximum channel capacity. Due to the cache-independent feature of PREFETCHX, current cache-based mitigations are not effective against our attacks. Overall, our work uncovers a significant vulnerability in the XPT prefetcher, which can be exploited to compromise the confidentiality of sensitive information in both cryptography and non-cryptography-related applications among processor cores.
Yun Chen 0004, Ali Hajiabadi, Lingfeng Pei, Trevor E. Carlson
HPCA2
2024 Efficient Detection and Mitigation Schemes for Speculative Side Channels
abstract
The introduction of Spectre in 2018 demonstrated a serious threat in almost all modern processors since Spectre exploits the main performance enabler of processors: speculative execution. Detecting and mitigating speculative execution attacks have been a major line of research in the past years. In this work, we explore new ways to bypass existing detection mechanisms and then propose a mitigation strategy to comprehensively prevent speculative execution vulnerabilities through the cache side-channel. Our results show that our proposed protection incurs almost zero performance overhead while improving security.
Arash Pashrashid, Ali Hajiabadi, Trevor E. Carlson
ISCAS2
2023 HidFix: Efficient Mitigation of Cache-Based Spectre Attacks Through Hidden Rollbacks
abstract
Mitigating Spectre attacks in modern systems is a challenging task for CPU vendors as they need to provide comprehensive protection while maintaining high efficiency. One common solution is to adopt always-on mitigation strategies to prevent all speculative data leaks. However, these solutions incur prohibitive performance overheads as they limit the benefits of speculative execution, the main performance enabler of modern processors. Additionally, recent attacks have demonstrated the limitations of many existing defenses. Combining side-channel attack (SCA) detectors with mitigation strategies is a promising direction to achieve efficient and selective mitigation of Spectre attacks. In this work, we enumerate the combinations of state-of-the-art detection and mitigation strategies and present both new attacks as well as the potential risks of such detection/mitigation combinations. The result is the HIDFIX methodology, an efficient mitigation for cache-based Spectre attacks, that addresses the security limitations of prior work. We show that Hidfix has a near-zero performance overhead for all evaluated applications. Hidfix rollbacks the misspeculated data leaks in a timely manner, before an attacker has the chance to infer the victim's sensitive data. We demonstrate that HidFix is more secure compared to prior cache-based Spectre defenses, and moreover, it does not introduce new side effects that might enable an attacker to observe secret dependent changes in the system.
Arash Pashrashid, Ali Hajiabadi, Trevor E. Carlson
ICCAD2
2022 Fast, Robust and Accurate Detection of Cache-Based Spectre Attack Phases
abstract
Modern processors achieve high performance and efficiency by employing techniques such as speculative execution and sharing resources such as caches. However, recent attacks like Spectre and Meltdown exploit the speculative execution of modern processors to leak sensitive information from the system. Many mitigation strategies have been proposed to restrict the speculative execution of processors and protect potential side-channels. Currently, these techniques have shown a significant performance overhead. A solution that can detect memory leaks before the attacker has a chance to exploit them would allow the processor to reduce the performance overhead by enabling protections only when the system is at risk.
Arash Pashrashid, Ali Hajiabadi, Trevor E. Carlson
ICCAD2
2021 NOREBA: a compiler-informed non-speculative out-of-order commit processor
abstract
Modern superscalar processors execute instructions out-of-order, but commit them in program order to provide precise exception handling and safe instruction retirement. However, in-order instruction commit is highly conservative and holds on to critical resources far longer than necessary, severely limiting the reach of general-purpose processors, ultimately reducing performance. Solutions that allow for efficient, early reclamation of these critical resources could seize the opportunity to improve performance. One such solution is out-of-order commit, which has traditionally been challenging due to inefficient, complex hardware used to guarantee safe instruction retirement and provide precise exception handling.
Ali Hajiabadi, Andreas Diavastos, Trevor E. Carlson
ASPLOS1
2021 ELFies: Executable Region Checkpoints for Performance Analysis and Simulation
abstract
We address the challenge faced in characterizing long-running workloads, namely how to reliably focus the detailed analysis on interesting execution regions. We present a set of tools that allows users to precisely capture any region of interest in program execution, and create a stand-alone executable, called an ELFie, from it. An ELFie starts with the same program state captured at the beginning of the region of interest and then executes natively. With ELFies, there is no fast-forwarding to the region of interest needed or the uncertainty of reaching the region. ELFies can be fed to dynamic program-analysis tools or simulators that work with regular program binaries. Our tool-chain is based on the PinPlay framework and requires no special hardware, operating system changes, recompilation, or re-linking of test programs. This paper describes the design of our ELFie generation tool-chain and the application of ELFies in performance analysis and simulation of regions of interest in popular long-running single and multi-threaded benchmarks.
Harish Patil, Alexander Isaev, Wim Heirman, Alen Sabu, Ali Hajiabadi, Trevor E. Carlson
CGO5
2019 Highly Concurrent Latency-tolerant Register Files for GPUs
abstract
Graphics Processing Units (GPUs) employ large register files to accommodate all active threads and accelerate context switching. Unfortunately, register files are a scalability bottleneck for future GPUs due to long access latency, high power consumption, and large silicon area provisioning. Prior work proposes hierarchical register file to reduce the register file power consumption by caching registers in a smaller register file cache. Unfortunately, this approach does not improve register access latency due to the low hit rate in the register file cache. In this article, we propose the Latency-Tolerant Register File (LTRF) architecture to achieve low latency in a two-level hierarchical structure while keeping power consumption low. We observe that compile-time interval analysis enables us to divide GPU program execution into intervals with an accurate estimate of a warp’s aggregate register working-set within each interval. The key idea of LTRF is to prefetch the estimated register working-set from the main register file to the register file cache under software control, at the beginning of each interval, and overlap the prefetch latency with the execution of other warps. We observe that register bank conflicts while prefetching the registers could greatly reduce the effectiveness of LTRF. Therefore, we devise a compile-time register renumbering technique to reduce the likelihood of register bank conflicts. Our experimental results show that LTRF enables high-capacity yet long-latency main GPU register files, paving the way for various optimizations. As an example optimization, we implement the main register file with emerging high-density high-latency memory technologies, enabling 8× larger capacity and improving overall GPU performance by 34%.
Mohammad Sadrosadati, Amirhossein Mirhosseini, Ali Hajiabadi, Seyed Borna Ehsani, Hajar Falahati, Hamid Sarbazi-Azad, Mario Drumond, Babak Falsafi, Rachata Ausavarungnirun, Onur Mutlu
ACM Trans. Comput. Syst.3