VLDB 2026 Research / reviewers in the wild / expert
Mehdi Alipour
dblp:41/9546
· DBLP profile ↗
10ranked-venue papers
4as first author
2since 2021 · last 2026
0000-0001-9842-8715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 62% Memory systems · 24% Electronic design automation · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Concurrent programming · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › register file
physical register file |
1.0 | 1 | 2026 | Tempranillo: Non-Speculative Early Register Release · HPCA 2026 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
1.0 | 1 | 2026 | Tempranillo: Non-Speculative Early Register Release · HPCA 2026 |
Processor architecture and microarchitecture
register file |
1.0 | 1 | 2026 | Tempranillo: Non-Speculative Early Register Release · HPCA 2026 |
Memory systems › memory access optimization
memory-level parallelism |
1.0 | 2 | 2022 | Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order Cores · ACM Trans. Archit. Code Optim. 2022 Freeway: Maximizing MLP for Slice-Out-of-Order Execution · HPCA 2019 |
Processor architecture and microarchitecture
out-of-order execution |
1.0 | 2 | 2022 | Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order Cores · ACM Trans. Archit. Code Optim. 2022 Freeway: Maximizing MLP for Slice-Out-of-Order Execution · HPCA 2019 |
Processor architecture and microarchitecture › front-end
instruction queue |
0.4 | 1 | 2020 | Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020 |
Processor architecture and microarchitecture
instruction scheduling |
0.4 | 1 | 2020 | Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020 |
Processor architecture and microarchitecture › out-of-order execution
out-of-order scheduling |
0.4 | 1 | 2020 | Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020 |
Memory systems › memory consistency
memory consistency model |
0.4 | 2 | 2018 | Constructing a Weak Memory Model · ISCA 2018 Non-Speculative Load-Load Reordering in TSO · ISCA 2017 |
Concurrent programming
memory models |
0.3 | 1 | 2018 | Constructing a Weak Memory Model · ISCA 2018 |
Concurrent programming › concurrency semantics
operational and axiomatic semantics |
0.3 | 1 | 2018 | Constructing a Weak Memory Model · ISCA 2018 |
Memory systems › memory consistency › memory consistency model
weak memory model |
0.3 | 1 | 2018 | Constructing a Weak Memory Model · ISCA 2018 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.3 | 1 | 2026 | Tempranillo: Non-Speculative Early Register Release · HPCA 2026 |
Memory systems
cache coherence |
0.3 | 1 | 2017 | Non-Speculative Load-Load Reordering in TSO · ISCA 2017 |
Energy-efficient computing
power management |
0.1 | 1 | 2020 | Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order Processors · HPCA 2020 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.1 | 1 | 2018 | Constructing a Weak Memory Model · ISCA 2018 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.7operational semantics · 0.7axiomatic semantics · 0.7performance simulation · 0.6cycle-accurate simulation · 0.4performance evaluation · 0.4non-speculative reordering · 0.3directory protocol modification · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tempranillo: Non-Speculative Early Register ReleaseabstractLimited by the breakdown of technology scaling, CPU architects are looking for creative solutions to deliver performance improvements while unable to traditionally scale up microarchitectural structures. One promising approach is engineering a microarchitecture that uses resources more efficiently, for example, by recycling them faster to relieve pressure on critical structures and making them look bigger than they are. The physical register file (PRF) is a key structure that faces severe area and power constraints that limit its (and the whole CPU's) scalability. Based on this observation, previous work proposed solutions to reduce the pressure on the PRF by reducing the time each register remains allocated. In this paper, we corroborate earlier findings that the early release of registers is a promising approach to reduce the occupancy of the PRF and we identify novel tight conditions for safely releasing registers. Based on this analysis, we design Tempranillo: an aggressive, non-speculative microarchitecture to release registers as early as possible without requiring additional recovery mechanisms. Tempranillo delivers up to 3.3 % and 11.8 % performance improvement over conventional release on singlethreaded and a 2 -way SMT CPUs, respectively. Additionally, Tempranillo requires modest storage overheads, translating into a performance improvement per KiB of storage of up to 2.6 % and 9.3 % for single-thread and 2-way SMT, respectively. Our evaluation shows that Tempranillo improves over both the state-of-the-art non-speculative and speculative proposals. Carlos Escuin, Paolo Salvatore Galfano, Davide B. Bartolini, Leeor Peled, Mehdi Alipour |
HPCA | 5 |
| 2022 | Dependence-aware Slice Execution to Boost MLP in Slice-out-of-order CoresabstractExploiting memory-level parallelism (MLP) is crucial to hide long memory and last-level cache access latencies. While out-of-order (OoO) cores, and techniques building on them, are effective at exploiting MLP, they deliver poor energy efficiency due to their complex and energy-hungry hardware. This work revisits slice-out-of-order (sOoO) cores as an energy-efficient alternative for MLP exploitation. sOoO cores achieve energy efficiency by constructing and executing slices of MLP-generating instructions out-of-order only with respect to the rest of instructions; the slices and the remaining instructions, by themselves, execute in-order. However, we observe that existing sOoO cores miss significant MLP opportunities due to their dependence-oblivious in-order slice execution, which causes dependent slices to frequently block MLP generation. To boost MLP generation, we introduce Freeway, a sOoO core based on a new dependence-aware slice execution policy that tracks dependent slices and keeps them from blocking subsequent independent slices and MLP extraction. The proposed core incurs minimal area and power overheads, yet approaches the MLP benefits of fully OoO cores. Our evaluation shows that Freeway delivers 12% better performance than the state-of-the-art sOoO core and is within 7% of the MLP limits of full OoO execution. Rakesh Kumar 0003, Mehdi Alipour, David Black-Schaffer |
ACM Trans. Archit. Code Optim. | 2 |
| 2020 | Delay and Bypass: Ready and Criticality Aware Instruction Scheduling in Out-of-Order ProcessorsabstractFlexible instruction scheduling is essential for performance in out-of-order processors. This is typically achieved by using CAM-based Instruction Queues (IQs) that provide complete flexibility in choosing ready instructions for execution, but at the cost of significant scheduling energy. In this work we seek to reduce the instruction scheduling energy by reducing the depth and width of the IQ. We do so by classifying instructions based on their readiness and criticality, and using this information to bypass the IQ for instructions that will not benefit from its expensive scheduling structures and delay instructions that will not harm performance. Combined, these approaches allow us to offload a significant portion of the instructions from the IQ to much cheaper FIFO-based scheduling structures without hurting performance. As a result we can reduce the IQ depth and width by half, thereby saving energy. Our design, Delay and Bypass (DNB), is the first design to explicitly address both readiness and criticality to reduce scheduling energy. By handling both classes we are able to achieve 95% of the baseline out-of-order performance while only using 33% of the scheduling energy. This represents a significant improvement over previous designs which addressed only criticality or readiness (91%/89% performance at 74%/53% energy). Mehdi Alipour, Stefanos Kaxiras, David Black-Schaffer, Rakesh Kumar 0003 |
HPCA | 1 |
| 2019 | Ghost loads: what is the cost of invisible speculation?abstractSpeculative execution is necessary for achieving high performance on modern general-purpose CPUs but, starting with Spectre and Meltdown, it has also been proven to cause severe security flaws. In case of a misspeculation, the architectural state is restored to assure functional correctness but a multitude of microarchitectural changes (e.g., cache updates), caused by the speculatively executed instructions, are commonly left in the system. These changes can be used to leak sensitive information, which has led to a frantic search for solutions that can eliminate such security flaws. The contribution of this work is an evaluation of the cost of hiding speculative side-effects in the cache hierarchy, making them visible only after the speculation has been resolved. For this, we compare (for the first time) two broad approaches: i) waiting for loads to become non-speculative before issuing them to the memory system, and ii) eliminating the side-effects of speculation, a solution consisting of invisible loads (Ghost loads) and performance optimizations (Ghost Buffer and Materialization). While previous work, InvisiSpec, has proposed a similar solution to our latter approach, it has done so with only a minimal evaluation and at a significant performance cost. The detailed evaluation of our solutions shows that: i) waiting for loads to become non-speculative is no more costly than the previously proposed InvisiSpec solution, albeit much simpler, non-invasive in the memory system, and stronger security-wise; ii) hiding speculation with Ghost loads (in the context of a relaxed memory model) can be achieved at the cost of 12% performance degradation and 9% energy increase, which is significantly better that the previous state-of-the-art solution. Christos Sakalis, Mehdi Alipour, Alberto Ros 0001, Alexandra Jimborean, Stefanos Kaxiras, Magnus Själander |
CF | 2 |
| 2019 | FIFOrder MicroArchitecture: Ready-Aware Instruction Scheduling for OoO ProcessorsabstractThe number of instructions a processor's instruction queue can examine (depth) and the number it can issue together (width) determine its ability to take advantage of the ILP in an application. Unfortunately, increasing either the width or depth of the instruction queue is very costly due to the content-addressable logic needed to wakeup and select instructions out-of-order. This work makes the observation that a large number of instructions have both operands ready at dispatch, and therefore do not benefit from out-of-order scheduling. We leverage this to place such ready-at-dispatch instructions in separate, simpler, in-order FIFO queues for scheduling. With such additional queues, we can reduce the size and width of the expensive out-of-order instruction queue, without reducing the processor's overall issue width and depth. Our design, FIFOrder, is able to steer more than 60% of instructions to the cheaper FIFO queues, providing a 50% energy savings over a traditional out-of-order instruction queue design, while delivering 8% higher performance. Mehdi Alipour, Rakesh Kumar 0003, Stefanos Kaxiras, David Black-Schaffer |
DATE | 1 |
| 2019 | Freeway: Maximizing MLP for Slice-Out-of-Order ExecutionabstractExploiting memory level parallelism (MLP) is crucial to hide long memory and last level cache access latencies. While out-of-order (OoO) cores, and techniques building on them, are effective at exploiting MLP, they deliver poor energy efficiency due to their complex hardware and the resulting energy overheads. As energy efficiency becomes the prime design constraint, we investigate low complexity/energy mechanisms to exploit MLP. This work revisits slice-out-of-order (sOoO) cores as an energy efficient alternative to OoO cores for MLP exploitation. These cores construct slices of MLP generating instructions and execute them out-of-order with respect to the rest of instructions. However, the slices and the remaining instructions, by themselves, execute in-order. Though their energy overhead is low compared to full OoO cores, sOoO cores fall considerably behind in terms of MLP extraction. We observe that their dependence-oblivious inorder slice execution causes dependent slices to frequently block MLP generation. To boost MLP generation in sOoO cores, we introduce Freeway, a sOoO core based on a new dependence-aware slice execution policy that tracks dependent slices and keeps them out of the way of MLP extraction. The proposed core incurs minimal area and power overheads, yet approaches the MLP benefits of fully OoO cores. Our evaluation shows that Freeway outperforms the state-of-the-art sOoO core by 12% and is within 7% of the MLP limits of full OoO execution. Rakesh Kumar 0003, Mehdi Alipour, David Black-Schaffer |
HPCA | 2 |
| 2018 | Constructing a Weak Memory ModelabstractWeak memory models are a consequence of the desire on part of architects to preserve all the uniprocessor optimizations while building a shared memory multiprocessor. The efforts to formalize weak memory models of ARM and POWER over the last decades are mostly empirical – they try to capture empirically observed behaviors – and end up providing no insight into the inherent nature of weak memory models. This paper takes a constructive approach to find a common base for weak memory models: we explore what a weak memory would look like if we constructed it with the explicit goal of preserving all the uniprocessor optimizations. We will disallow some optimizations which break a programmer's intuition in highly unexpected ways. The constructed model, which we call General Atomic Memory Model (GAM), allows all four load/store reorderings. We give the construction procedure of GAM, and provide insights which are used to define its operational and axiomatic semantics. Though no attempt is made to match GAM to any existing weak memory model, we show by simulation that GAM has comparable performance with other models. No deep knowledge of memory models is needed to read this paper. Sizhuo Zhang, Muralidaran Vijayaraghavan, Andrew Wright, Mehdi Alipour, Arvind 0001 |
ISCA | 4 |
| 2017 | Non-Speculative Load-Load Reordering in TSOabstractIn Total Store Order memory consistency (TSO), loads can be speculatively reordered to improve performance. If a load-load reordering is seen by other cores, speculative loads must be squashed and re-executed. In architectures with an unordered interconnection network and directory coherence, this has been the established view for decades. We show, for the first time, that it is not necessary to squash and re-execute speculatively reordered loads in TSO when their reordering is seen. Instead, the reordering can be hidden form other cores by the coherence protocol. The implication is that we can irrevocably bind speculative loads. This allows us to commit reordered loads out-of-order without having to wait (for the loads to become non-speculative) or without having to checkpoint committed state (and rollback if needed), just to ensure correctness in the rare case of some core seeing the reordering. We show that by exposing a reordering to the coherence layer and by appropriately modifying a typical directory protocol we can successfully hide load-load reordering without perceptible performance cost and without deadlock. Our solution is cost-effective and increases the performance of out-of-order commit by a sizable margin, compared to the base case where memory operations are not allowed to commit if the consistency model could be violated. Alberto Ros 0001, Trevor E. Carlson, Mehdi Alipour, Stefanos Kaxiras |
ISCA | 3 |
| 2017 | A taxonomy of out-of-order instruction commitabstractWhile in-order instruction commit has its advantages, such as providing precise interrupts and avoiding complications with the memory consistency model, it requires the core to hold on to resources (reorder buffer entries, load/store queue entries, registers) until they are released in program order. In contrast, out-of-order commit releases resources much earlier, yielding improved performance without the need for additional hardware resources. In this paper, we revisit out-of-order commit from a different perspective, not by proposing another hardware technique, but by introducing a taxonomy and evaluating three different micro-architectures that have this technique enabled. We show how smaller processors can benefit from simple out-oforder commit strategies, but that larger, aggressive cores require more aggressive strategies to improve performance. Mehdi Alipour, Trevor E. Carlson, Stefanos Kaxiras |
ISPASS | 1 |
| 2011 | Congestion and track usage improvement of large FPGAs using metro-on-FPGA methodologyabstractAsynchronous serial transceivers have been recently used for data multiplexing in large on-chip systems to alleviate the routing congestion and improve the routability. FPGAs have considerable potential for using the serial transmission but these links have not been exploited in FPGAs yet. In this paper, we present a new architecture corresponding with a routing algorithm to use the asynchronous wire multiplexing technique in FPGAs. Experimental results show that allocated routing tracks and routing congestion can be reduced considerably (9.37% and 9.03%, respectively) by using the asynchronous wire multiplexing without any performance degradation in cost of a little overhead in area and computation time (2% and 0.84%, respectively) Mehdi Alipour, Mohammad Haji Seyed Javadi, Ali Jahanian 0001 |
ACM Great Lakes Symposium on VLSI | 1 |