Mehdi Alipour

dblp:41/9546 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
2since 2021 · last 2026
0000-0001-9842-8715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Processor architecture and microarchitecture · 62% Memory systems · 24% Electronic design automation · 12%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › register file
physical register file
1.012026
Tempranillo: Non-Speculative Early Register Release · HPCA 2026
Electronic design automation › high-level synthesis › resource binding
register allocation
1.012026
Tempranillo: Non-Speculative Early Register Release · HPCA 2026
Processor architecture and microarchitecture
register file
1.012026
Tempranillo: Non-Speculative Early Register Release · HPCA 2026
Memory systems › memory access optimization
memory-level parallelism
1.022022
Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order Cores · ACM Trans. Archit. Code Optim. 2022
Freeway: Maximizing MLP for Slice-Out-of-Order Execution · HPCA 2019
Processor architecture and microarchitecture
out-of-order execution
1.022022
Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order Cores · ACM Trans. Archit. Code Optim. 2022
Freeway: Maximizing MLP for Slice-Out-of-Order Execution · HPCA 2019
Processor architecture and microarchitecture › front-end
instruction queue
0.412020
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020
Processor architecture and microarchitecture
instruction scheduling
0.412020
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020
Processor architecture and microarchitecture › out-of-order execution
out-of-order scheduling
0.412020
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020
Memory systems › memory consistency
memory consistency model
0.422018
Constructing a Weak Memory Model · ISCA 2018
Non-Speculative Load-Load Reordering in TSO · ISCA 2017
Concurrent programming
memory models
0.312018
Constructing a Weak Memory Model · ISCA 2018
Concurrent programming › concurrency semantics
operational and axiomatic semantics
0.312018
Constructing a Weak Memory Model · ISCA 2018
Memory systems › memory consistency › memory consistency model
weak memory model
0.312018
Constructing a Weak Memory Model · ISCA 2018
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.312026
Tempranillo: Non-Speculative Early Register Release · HPCA 2026
Memory systems
cache coherence
0.312017
Non-Speculative Load-Load Reordering in TSO · ISCA 2017
Energy-efficient computing
power management
0.112020
Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.112018
Constructing a Weak Memory Model · ISCA 2018

Methods — techniques the papers use, named apart from their topics

simulation · 0.7operational semantics · 0.7axiomatic semantics · 0.7performance simulation · 0.6cycle-accurate simulation · 0.4performance evaluation · 0.4non-speculative reordering · 0.3directory protocol modification · 0.3
YearPublicationVenuePosition
2026 Tempranillo: Non-Speculative Early Register Release
abstract
Limited by the breakdown of technology scaling, CPU architects are looking for creative solutions to deliver performance improvements while unable to traditionally scale up microarchitectural structures. One promising approach is engineering a microarchitecture that uses resources more efficiently, for example, by recycling them faster to relieve pressure on critical structures and making them look bigger than they are. The physical register file (PRF) is a key structure that faces severe area and power constraints that limit its (and the whole CPU's) scalability. Based on this observation, previous work proposed solutions to reduce the pressure on the PRF by reducing the time each register remains allocated. In this paper, we corroborate earlier findings that the early release of registers is a promising approach to reduce the occupancy of the PRF and we identify novel tight conditions for safely releasing registers. Based on this analysis, we design Tempranillo: an aggressive, non-speculative microarchitecture to release registers as early as possible without requiring additional recovery mechanisms. Tempranillo delivers up to 3.3 % and 11.8 % performance improvement over conventional release on singlethreaded and a 2 -way SMT CPUs, respectively. Additionally, Tempranillo requires modest storage overheads, translating into a performance improvement per KiB of storage of up to 2.6 % and 9.3 % for single-thread and 2-way SMT, respectively. Our evaluation shows that Tempranillo improves over both the state-of-the-art non-speculative and speculative proposals.
Carlos Escuin, Paolo Salvatore Galfano, Davide B. Bartolini, Leeor Peled, Mehdi Alipour
HPCA5
2022 Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order Cores
abstract
Exploiting memory-level parallelism (MLP) is crucial to hide long memory and last-level cache access latencies. While out-of-order (OoO) cores, and techniques building on them, are effective at exploiting MLP, they deliver poor energy efficiency due to their complex and energy-hungry hardware. This work revisits slice-out-of-order (sOoO) cores as an energy-efficient alternative for MLP exploitation. sOoO cores achieve energy efficiency by constructing and executing slices of MLP-generating instructions out-of-order only with respect to the rest of instructions; the slices and the remaining instructions, by themselves, execute in-order. However, we observe that existing sOoO cores miss significant MLP opportunities due to their dependence-oblivious in-order slice execution, which causes dependent slices to frequently block MLP generation. To boost MLP generation, we introduce Freeway, a sOoO core based on a new dependence-aware slice execution policy that tracks dependent slices and keeps them from blocking subsequent independent slices and MLP extraction. The proposed core incurs minimal area and power overheads, yet approaches the MLP benefits of fully OoO cores. Our evaluation shows that Freeway delivers 12% better performance than the state-of-the-art sOoO core and is within 7% of the MLP limits of full OoO execution.
Rakesh Kumar 0003, Mehdi Alipour, David Black-Schaffer
ACM Trans. Archit. Code Optim.2
2020 Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors
abstract
Flexible instruction scheduling is essential for performance in out-of-order processors. This is typically achieved by using CAM-based Instruction Queues (IQs) that provide complete flexibility in choosing ready instructions for execution, but at the cost of significant scheduling energy. In this work we seek to reduce the instruction scheduling energy by reducing the depth and width of the IQ. We do so by classifying instructions based on their readiness and criticality, and using this information to bypass the IQ for instructions that will not benefit from its expensive scheduling structures and delay instructions that will not harm performance. Combined, these approaches allow us to offload a significant portion of the instructions from the IQ to much cheaper FIFO-based scheduling structures without hurting performance. As a result we can reduce the IQ depth and width by half, thereby saving energy. Our design, Delay and Bypass (DNB), is the first design to explicitly address both readiness and criticality to reduce scheduling energy. By handling both classes we are able to achieve 95% of the baseline out-of-order performance while only using 33% of the scheduling energy. This represents a significant improvement over previous designs which addressed only criticality or readiness (91%/89% performance at 74%/53% energy).
Mehdi Alipour, Stefanos Kaxiras, David Black-Schaffer, Rakesh Kumar 0003
HPCA1
2019 Ghost loads: what is the cost of invisible speculation?
abstract
Speculative execution is necessary for achieving high performance on modern general-purpose CPUs but, starting with Spectre and Meltdown, it has also been proven to cause severe security flaws. In case of a misspeculation, the architectural state is restored to assure functional correctness but a multitude of microarchitectural changes (e.g., cache updates), caused by the speculatively executed instructions, are commonly left in the system. These changes can be used to leak sensitive information, which has led to a frantic search for solutions that can eliminate such security flaws. The contribution of this work is an evaluation of the cost of hiding speculative side-effects in the cache hierarchy, making them visible only after the speculation has been resolved. For this, we compare (for the first time) two broad approaches: i) waiting for loads to become non-speculative before issuing them to the memory system, and ii) eliminating the side-effects of speculation, a solution consisting of invisible loads (Ghost loads) and performance optimizations (Ghost Buffer and Materialization). While previous work, InvisiSpec, has proposed a similar solution to our latter approach, it has done so with only a minimal evaluation and at a significant performance cost. The detailed evaluation of our solutions shows that: i) waiting for loads to become non-speculative is no more costly than the previously proposed InvisiSpec solution, albeit much simpler, non-invasive in the memory system, and stronger security-wise; ii) hiding speculation with Ghost loads (in the context of a relaxed memory model) can be achieved at the cost of 12% performance degradation and 9% energy increase, which is significantly better that the previous state-of-the-art solution.
Christos Sakalis, Mehdi Alipour, Alberto Ros 0001, Alexandra Jimborean, Stefanos Kaxiras, Magnus Själander
CF2
2019 FIFOrder MicroArchitecture: Ready-Aware Instruction Scheduling for OoO Processors
abstract
The number of instructions a processor's instruction queue can examine (depth) and the number it can issue together (width) determine its ability to take advantage of the ILP in an application. Unfortunately, increasing either the width or depth of the instruction queue is very costly due to the content-addressable logic needed to wakeup and select instructions out-of-order. This work makes the observation that a large number of instructions have both operands ready at dispatch, and therefore do not benefit from out-of-order scheduling. We leverage this to place such ready-at-dispatch instructions in separate, simpler, in-order FIFO queues for scheduling. With such additional queues, we can reduce the size and width of the expensive out-of-order instruction queue, without reducing the processor's overall issue width and depth. Our design, FIFOrder, is able to steer more than 60% of instructions to the cheaper FIFO queues, providing a 50% energy savings over a traditional out-of-order instruction queue design, while delivering 8% higher performance.
Mehdi Alipour, Rakesh Kumar 0003, Stefanos Kaxiras, David Black-Schaffer
DATE1
2019 Freeway: Maximizing MLP for Slice-Out-of-Order Execution
abstract
Exploiting memory level parallelism (MLP) is crucial to hide long memory and last level cache access latencies. While out-of-order (OoO) cores, and techniques building on them, are effective at exploiting MLP, they deliver poor energy efficiency due to their complex hardware and the resulting energy overheads. As energy efficiency becomes the prime design constraint, we investigate low complexity/energy mechanisms to exploit MLP. This work revisits slice-out-of-order (sOoO) cores as an energy efficient alternative to OoO cores for MLP exploitation. These cores construct slices of MLP generating instructions and execute them out-of-order with respect to the rest of instructions. However, the slices and the remaining instructions, by themselves, execute in-order. Though their energy overhead is low compared to full OoO cores, sOoO cores fall considerably behind in terms of MLP extraction. We observe that their dependence-oblivious inorder slice execution causes dependent slices to frequently block MLP generation. To boost MLP generation in sOoO cores, we introduce Freeway, a sOoO core based on a new dependence-aware slice execution policy that tracks dependent slices and keeps them out of the way of MLP extraction. The proposed core incurs minimal area and power overheads, yet approaches the MLP benefits of fully OoO cores. Our evaluation shows that Freeway outperforms the state-of-the-art sOoO core by 12% and is within 7% of the MLP limits of full OoO execution.
Rakesh Kumar 0003, Mehdi Alipour, David Black-Schaffer
HPCA2
2018 Constructing a Weak Memory Model
abstract
Weak memory models are a consequence of the desire on part of architects to preserve all the uniprocessor optimizations while building a shared memory multiprocessor. The efforts to formalize weak memory models of ARM and POWER over the last decades are mostly empirical – they try to capture empirically observed behaviors – and end up providing no insight into the inherent nature of weak memory models. This paper takes a constructive approach to find a common base for weak memory models: we explore what a weak memory would look like if we constructed it with the explicit goal of preserving all the uniprocessor optimizations. We will disallow some optimizations which break a programmer's intuition in highly unexpected ways. The constructed model, which we call General Atomic Memory Model (GAM), allows all four load/store reorderings. We give the construction procedure of GAM, and provide insights which are used to define its operational and axiomatic semantics. Though no attempt is made to match GAM to any existing weak memory model, we show by simulation that GAM has comparable performance with other models. No deep knowledge of memory models is needed to read this paper.
Sizhuo Zhang, Muralidaran Vijayaraghavan, Andrew Wright, Mehdi Alipour, Arvind 0001
ISCA4
2017 Non-Speculative Load-Load Reordering in TSO
abstract
In Total Store Order memory consistency (TSO), loads can be speculatively reordered to improve performance. If a load-load reordering is seen by other cores, speculative loads must be squashed and re-executed. In architectures with an unordered interconnection network and directory coherence, this has been the established view for decades. We show, for the first time, that it is not necessary to squash and re-execute speculatively reordered loads in TSO when their reordering is seen. Instead, the reordering can be hidden form other cores by the coherence protocol. The implication is that we can irrevocably bind speculative loads. This allows us to commit reordered loads out-of-order without having to wait (for the loads to become non-speculative) or without having to checkpoint committed state (and rollback if needed), just to ensure correctness in the rare case of some core seeing the reordering. We show that by exposing a reordering to the coherence layer and by appropriately modifying a typical directory protocol we can successfully hide load-load reordering without perceptible performance cost and without deadlock. Our solution is cost-effective and increases the performance of out-of-order commit by a sizable margin, compared to the base case where memory operations are not allowed to commit if the consistency model could be violated.
Alberto Ros 0001, Trevor E. Carlson, Mehdi Alipour, Stefanos Kaxiras
ISCA3
2017 A taxonomy of out-of-order instruction commit
abstract
While in-order instruction commit has its advantages, such as providing precise interrupts and avoiding complications with the memory consistency model, it requires the core to hold on to resources (reorder buffer entries, load/store queue entries, registers) until they are released in program order. In contrast, out-of-order commit releases resources much earlier, yielding improved performance without the need for additional hardware resources. In this paper, we revisit out-of-order commit from a different perspective, not by proposing another hardware technique, but by introducing a taxonomy and evaluating three different micro-architectures that have this technique enabled. We show how smaller processors can benefit from simple out-oforder commit strategies, but that larger, aggressive cores require more aggressive strategies to improve performance.
Mehdi Alipour, Trevor E. Carlson, Stefanos Kaxiras
ISPASS1
2011 Congestion and track usage improvement of large FPGAs using metro-on-FPGA methodology
abstract
Asynchronous serial transceivers have been recently used for data multiplexing in large on-chip systems to alleviate the routing congestion and improve the routability. FPGAs have considerable potential for using the serial transmission but these links have not been exploited in FPGAs yet. In this paper, we present a new architecture corresponding with a routing algorithm to use the asynchronous wire multiplexing technique in FPGAs. Experimental results show that allocated routing tracks and routing congestion can be reduced considerably (9.37% and 9.03%, respectively) by using the asynchronous wire multiplexing without any performance degradation in cost of a little overhead in area and computation time (2% and 0.84%, respectively)
Mehdi Alipour, Mohammad Haji Seyed Javadi, Ali Jahanian 0001
ACM Great Lakes Symposium on VLSI1