Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ashkan Asgharzadeh

dblp:321/3713 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 42% Parallel and multicore computing · 32% Processor architecture and microarchitecture · 26%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
memory consistency
0.912025
No Rush in Executing Atomic Instructions · HPCA 2025
Parallel and multicore computing
synchronization
0.912025
No Rush in Executing Atomic Instructions · HPCA 2025
Memory systems › memory consistency
x86-TSO
0.912025
No Rush in Executing Atomic Instructions · HPCA 2025
Memory systems › memory consistency
memory fences
0.612022
Free atomics: hardware atomic operations without fences · ISCA 2022
Parallel and multicore computing › synchronization
read-modify-write
0.612022
Free atomics: hardware atomic operations without fences · ISCA 2022
Processor architecture and microarchitecture › load/store queue
store buffer
0.612022
Free atomics: hardware atomic operations without fences · ISCA 2022
Parallel and multicore computing › synchronization
synchronization mechanisms
0.612022
Free atomics: hardware atomic operations without fences · ISCA 2022
Memory systems › memory interference
cache line contention
0.312025
No Rush in Executing Atomic Instructions · HPCA 2025
Processor architecture and microarchitecture › pipelining
pipeline design
0.212022
Free atomics: hardware atomic operations without fences · ISCA 2022

Methods — techniques the papers use, named apart from their topics

simulation · 0.9contention prediction · 0.9
YearPublicationVenuePosition
2025 No Rush in Executing Atomic Instructions
abstract
Hardware atomic instructions are the building blocks of the synchronization algorithms. Historically, to guarantee atomicity and consistency, they were implemented using memory fences, committing older memory instructions, and draining the store buffer before initiating the execution of atomics. Unfortunately, the use of such memory fences entails huge performance penalties as it implies execution serialization, thus impeding instruction- and memory-level parallelism. The situation, however, seems to have changed recently. Through experiments on $x 86$ machines, we discovered that current $x 86$ processors manage to comply with the x86-TSO requirements while avoiding the performance overhead introduced by fences (fence-free or unfenced implementation). This paves the way to new potential optimizations to atomic instruction execution. In particular, our simulation experiments modeling unfenced atomics reveal that executing atomic instructions as soon as their operands are ready does not always lead to optimal performance. In fact, this increases the time that other threads should wait to obtain the cacheline. In contended scenarios, delaying the execution of the atomic instruction to minimize the time the cacheline is locked provides superior performance. Based on this observation, we present Rush or Wait (RoW), a hardware mechanism to decide when to execute an atomic instruction. The mechanism is based on a contention predictor that estimates if an atomic will access a contended cacheline. Non-contended atomics execute once their operands are ready. Contended atomics, on the contrary, wait to become the oldest memory instruction and to drain the store buffer to execute, minimizing the contention on the accessed cacheline. Our experimental evaluation shows that RoW reduces execution time on average by 9.2% (and up to 43%) compared to a baseline that executes atomics as soon as the operands are ready, and yet it requires a small area overhead (64 bytes).
Ashkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras, Alberto Ros 0001
HPCA1
2024 Hardware Cache Locking for All Memory Updates
abstract
Many applications need to perform operations that involve reading a value from memory, modifying it, and then writing it back. Multiple architectures provide hardware support for these operations via read-modify-write (RMW) instructions. The primary benefit is that the read can request a cacheline with write permissions, reducing coherence protocol overhead since the write will find the cacheline with appropriate permissions. RMWs can be either atomic or non-atomic. Atomic RMWs, used for synchronization, commonly require (i) locking the cacheline to guarantee atomicity by preventing invalidations and (ii) enforcing serialization of instructions in the program (e.g., via memory fences), which may cause performance degradation based on the implemented memory consistency model. Non-atomic RMWs, while not requiring such strict measures, should only be used in data-race free code sections. However, other cores may invalidate a cacheline during a non-atomic RMW (e.g., due to false sharing), flushing the pipeline and causing the loss of write permissions obtained by the read, which is detrimental to performance. In this work, we propose a microarchitectural mechanism that enables non-atomic RMWs to fetch the cacheline locking it, thus preventing other cores from “stealing” the cacheline while allowing them to run concurrently with other instructions in the same core. Our proposal enables concurrent hardware cache locking for multiple non-atomic RMW s while guaranteeing deadlock freedom and no programmer/compiler intervention. We also propose a lock-chaining mechanism to allow multiple consecutive memory updates to the same cacheline up to a predefined maximum (to prevent starvation and load imbalance). Our evaluation using gem5 full-system simulator shows that for an eight-core configuration, our proposal improves performance by up to 5.36 % (2.05 % on average), requiring just 45 bytes of storage per core.
Ashkan Asgharzadeh, Eduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto Ros 0001
ICCD1
2022 Free atomics: hardware atomic operations without fences
abstract
Atomic Read-Modify-Write (RMW) instructions are primitive synchronization operations implemented in hardware that provide the building blocks for higher-abstraction synchronization mechanisms to programmers. According to publicly available documentation, current x86 implementations serialize atomic RMW operations, i.e., the store buffer is drained before issuing atomic RMWs and subsequent memory operations are stalled until the atomic RMW commits. This serialization, carried out by memory fences, incurs a performance cost which is expected to increase with deeper pipelines.
Ashkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras, Alberto Ros 0001
ISCA1