VLDB 2026 Research / reviewers in the wild / expert
Rashid Aligholipour
dblp:264/8107
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-9376-2925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NoCWalk: In-Network Page Walks for Efficient Pointer-Chasing Workloads on Multicores
Yuan Yao 0009, Rashid Aligholipour, Stefanos Kaxiras |
CF | 3 |
| 2026 | DICE: Detailed Inter-Chiplet End-to-End Phy Modeling for Accurate Chiplet Simulation
Rashid Aligholipour, Stefanos Kaxiras, Yuan Yao 0009 |
ISCA | 1 |
| 2026 | Understanding Simulated Architecture via gem5 Call-Stack ProfilingabstractUnderstanding the behavior of simulated architectures in gem5 is critical for studying complex, deeply integrated computing systems. However, conventional analysis methods, which rely heavily on simulation statistics, provide only an indirect view of the simulated system internals. In this work, we show that call-stack profiling of gem5 itself offers a powerful yet underutilized perspective: the simulator’s own call-stack directly reflects the activity of the simulated system, exposing insights that conventional statistics may overlook.Profiling gem5’s call-stacks, however, is challenging due to its highly layered and complex software design patterns. To address this, we introduce a specialized, lightweight profiling framework built on Linux’s perf_event interface which samples and analyzes gem5’s runtime call-stacks throughout the simulation, resolves symbols on the fly, and merges samples into a hierarchical call-tree representation supporting both high-level structural views and focused, user-defined, component-specific analysis. Moreover, all profiling is performed in a dedicated helper process running alongside the main gem5 process, avoiding intrusive changes and overheads to the simulation itself.We apply our framework to gem5’s three major CPU modelsAtomicSimpleCPU, TimingSimpleCPU, and O3CPU-together with the Ruby memory system, and uncover behaviors that are not easily observable in conventional gem5 statistics. Our case studies reveal, for example, that TimingSimpleCPU is inefficient due to its use of a lockup-cache model and, despite its conceptual simplicity, does not simulate faster than a full out-of-order core. In addition, our tool makes it straightforward to detect cache coherence protocol deadlock and livelock-issues that are otherwise difficult to identify, since the simulation either appears to run normally or terminates abruptly, making it hard to pinpoint when these conditions occur. Johan Söderström, Rashid Aligholipour, Yuan Yao 0009 |
ISPASS | 2 |
| 2025 | RXT: RefleXive Address Translation for Pointer-Chasing WorkloadsabstractWith increasingly irregular memory access patterns in Indirect Memory Access (IMA) and Graph Processing (GP) applications, Virtual Address Translation (VAT) not only has become a major performance bottleneck, but also a key contributor to high power consumption in out-of-order (OoO) multi-core processor pipeline. We find that this problem can be attributed to a large extent to an inefficient utilization of TLBs in pointer-chasing instructions (i.e. load-to-load), which constitute a significant portion of workloads in IMA/GP. Conventionally, VATs for pointer-chasing are performed via a series of TLB look-ups: one for each load in the pointer chain. However, this approach is sub-optimal as it overlooks the correlations between pointers, missing opportunities to perform VATs through more costeffective calculations instead of relying on the more expensive TLB look-ups. This work draws on the insight that in pointer-chasing workloads, the physical page holding an upstream pointer is often at a fixed distance to the physical page where a downstream pointer resides. Building on this insight, we introduce RefleXive address Translation (RXT), which encodes the physical page distance (termed PageDist) between an upstream and downstream pointer into the unused upper 16 bits of the upstream pointer's virtual address. Consequently, RXT can compute the translation of a pointer by directly adding the PageDist to its residing address (the physical address where the pointer is stored). This process can be recursively applied throughout the pointer-chasing sequence, transforming address translation from a sequence of TLB accesses to computes. Through gem5 full-system simulation, RXT reduces TLB energy consumption for pointer-chasing applications by an average of 41.48 % (up to 62.50 %) and decreases overall core power density by 4.65 % (up to 13.09 %), without compromising application execution time. In fact, RXT even improves program runtime by an average of 0.71 % (up to 2.09 %). Rashid Aligholipour, Pavlos Aimoniotis, Stefanos Kaxiras, Yuan Yao 0009 |
IPDPS | 1 |
| 2025 | The Fake-Busy and True-Idle Problems of Running Graph Applications on Chiplet-Based Multi-CoresabstractWe introduce the fake-busy and true-idle problems encountered when running large graph workloads on chipletbased Out-of-Order (OoO) multi-cores. Caused by high interchiplet communication latency and irregular memory access patterns, these issues lead to inefficient use of core pipeline resources such as the Reorder Buffer (RoB) and Load Queue (LQ). Our evaluation shows that reducing RoB and LQ sizes has minimal impact on application performance, revealing new performance optimization opportunities for large graph workloads. Rashid Aligholipour, Yuan Yao 0009 |
ISPASS | 1 |
| 2021 | TAMA: Turn-aware Mapping and Architecture - A Power-efficient Network-on-Chip ApproachabstractNowadays, static power consumption in chip multiprocessor (CMP) is the most crucial concern of chip designers. Power-gating is an effective approach to mitigate static power consumption particularly in low utilization. Network-on-Chip (NoC) as the backbone of multi- and many-core chips has no exception. Previous state-of-the-art techniques in power-gating desire to decrease static power consumption alongside the lack of diminution in performance of NoC. However, maintaining the performance and utilization of the power-gating approach has not yet been addressed very well. In this article, we propose TAMA (Turn-Aware Mapping & Architecture) as an effective method to boost the performance of the TooT method that was only powering on a router during turning pass or packet injection. In other words, in the TooT method, straight and eject packets pass the router via a bypass route without powering on the router. By employing meta-heuristic approaches (Genetic and Ant Colony algorithms), we develop a specific application mapping that attempts to decrease the number of turns through interconnection networks. Accordingly, the average latency of packet transmission decreases due to fewer turns. Also, by powering on turn routers in advance with lightweight hardware, the latency of sending packets diminishes. The experimental results demonstrate that our proposed approach, i.e., TAMA achieves more than 13% reduction in packet latency of NoC in comparison with TooT. Besides the packet latency, the power consumption of TAMA is reduced by about 87% compared to the traditional approach. Rashid Aligholipour, Mohammad Baharloo, Behnam Farzaneh, Meisam Abdollahi, Ahmad Khonsari |
ACM Trans. Embed. Comput. Syst. | 1 |