VLDB 2026 Research / reviewers in the wild / expert
Ayaz Akram
dblp:190/5161
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-2201-6551ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Caching with A Tag-enhanced DRAMabstractAs SRAM-based caches are hitting a scaling wall, manufacturers are integrating DRAM-based caches into system designs to continue increasing cache sizes. While DRAM caches can improve the performance of memory systems, existing DRAM cache designs suffer from high miss penalties, wasted data movement, and interference between misses and demands. In this paper, we propose TDRAM, a novel DRAM microarchitecture tailored for caching. TDRAM enhances existing DRAM, such as HBM3, by adding small, low-latency mats to store tags and metadata on the same die as the data mats. These mats enable tag and data access in lockstep, in-DRAM tag comparison, and conditional data response based on the comparison result (reducing wasted data transfers), akin to SRAM cache mechanisms. TDRAM further optimizes hit and miss latencies through opportunistic early tag probing. Moreover, TDRAM introduces a flush buffer to store conflicting dirty data on write misses, eliminating data bus turnaround delays on write demands. We evaluate TDRAM in a full-system simulation using a set of HPC workloads with large memory footprints, showing that TDRAM, on average, provides $2.65 \times$ faster tag checks, $1.23 \times$ speedup, and 21% less energy consumption compared to state-of-the-art commercial and research designs. Maryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness, Michael R. Miller, Taeksang Song, Thomas Vogelsang, Steven C. Woo, Jason Lowe-Power |
HPCA | 2 |
| 2023 | Enabling Design Space Exploration of DRAM Caches for Emerging Memory SystemsabstractThe increasing growth of applications’ memory capacity and performance demands has led the CPU vendors to deploy heterogeneous memory systems either within a single system or via disaggregation. DRAM caches are one way to enable heterogeneity and disaggregation in such systems. While there is significant research investigating the designs of DRAM caches, there has been little research investigating DRAM caches from a full system point of view, because there is not a suitable model available to the community to accurately study large-scale systems with DRAM caches at a cycle-level. In this work we describe a new cycle-level DRAM cache model in the gem5 simulator which can be used for emerging heterogeneous and disaggregated memory systems. Maryam Babaie, Ayaz Akram, Jason Lowe-Power |
ISPASS | 2 |
| 2021 | Performance Analysis of Scientific Computing Workloads on General Purpose TEEsabstractScientific computing sometimes involves computation on sensitive data. Depending on the data and the execution environment, the HPC (high-performance computing) user or data provider may require confidentiality and/or integrity guarantees. To study the applicability of hardware-based trusted execution environments (TEEs) to enable secure scientific computing, we deeply analyze the performance impact of general purpose TEEs, AMD SEV, and Intel SGX, for diverse HPC benchmarks including traditional scientific computing, machine learning, graph analytics, and emerging scientific computing workloads. We observe three main findings: 1) SEV requires careful memory placement on large scale NUMA machines (1×-3.4× slowdown without and 1×-1.15× slowdown with NUMA aware placement), 2) virtualization-a prerequisite for SEV- results in performance degradation for workloads with irregular memory accesses and large working sets (1×-4× slowdown compared to native execution for graph applications) and 3) SGX is inappropriate for HPC given its limited secure memory size and inflexible programming model (1.2×-126× slowdown over unsecure execution). Finally, we discuss forthcoming new TEE designs and their potential impact on scientific computing. Ayaz Akram, Anna Giannakou, Venkatesh Akella, Jason Lowe-Power, Sean Peisert |
IPDPS | 1 |
| 2021 | Enabling Reproducible and Agile Full-System SimulationabstractRunning experiments in modern computer architecture simulators can be a difficult and error-prone endeavor. Users must track many configurations, components and outputs between simulation runs. The gem5 simulator is no exception to this, requiring researchers to gather, organize, and create a significant number of components for a single simulation. In this paper, we present the gem5art framework, a tool to aid gem5 users in better structuring and running architecture simulations, and gem5 resources, a suite of resources with known compatibility with the simulator. These new additions to the gem5 project make full system simulation easier, allowing researchers to concentrate more so on their architectural innovations over setting up the simulation framework. The gem5art framework carefully logs the resources used in a gem5 simulation and places the results obtained within a database, thus enabling simple reproduction of experiments. The pre-built resources allow researchers to jump straight into running simulations rather than having to spend valuable time creating them. gem5art has been released with a permissive, open source license allowing the broader computer architecture community to contribute as workloads and workflows evolve. An archive of the data, an related materials, presented in this paper can be found at https://doi.org/10.6084/m9.figshare.14176802. Bobby R. Bruce, Ayaz Akram, Hoa Nguyen, Kyle Roarty, Mahyar Samani, Marjan Fariborz, Trivikram Reddy, Matthew D. Sinclair, Jason Lowe-Power |
ISPASS | 2 |
| 2019 | A Study of Performance and Power Consumption Differences Among Different ISAsabstractRecent advances in different instruction set architectures (ISAs) and their implementations have revived the argument on the role of ISAs in the overall performance and energy efficiency of a processor. Many computer architects believe that with current compiler and microarchitecture developments, the choice of ISA is not a decisive matter anymore. On the other hand, some believe that ISAs can still play an important role in the overall performance and energy efficiency of a computer system. Our objective is to compare and contrast ISAs by finding the differences in performance and energy consumption across ISAs and the reasons behind those differences. Our work shows that ISAs affect the performance and energy efficiency of applications differently based on their inherent characteristics. Ayaz Akram, Lina Sawalha |
DSD | 1 |
| 2019 | FlexCPU: A Configurable Out-of-Order CPU AbstractionabstractWe present FlexCPU, a new software model for CPU performance integrated into gem5. FlexCPU combines the benefits of trace-based models with execute-in-execute semantics which leads to more accurate simulation of multithreaded and full-system applications. Our design is heavily inspired by dataflow models, and it reduces modern out-of-order techniques to abstracted parameterized constraints. FlexCPU can be configured to match the behaviors of modern general purpose CPUs and used for limit studies. By reducing CPU behaviors to reasonable abstractions and stages, FlexCPU is simpler to understand and easier to extend than other execute-in-execute CPU models. We show that FlexCPU can achieve the maximum theoretical ILP for most workloads and show a case study of using FlexCPU to model multiple processor architectures. Bradley Wang, Ayaz Akram, Jason Lowe-Power |
ISPASS | 2 |
| 2016 | ×86 computer architecture simulators: A comparative studyabstractThe significance of computer architecture simulators in advancing computer architecture research is widely acknowledged. Computer architects have developed numerous simulators in the past few decades and their number continues to rise. This paper explores different simulation techniques and surveys many ×86 simulators. Comparing simulators with each other and validating their correctness has been a challenging task. In this paper, we compare and contrast ×86 simulators in terms of flexibility, level of details, user friendliness and simulation models. In addition, we measure the experimental error and compare the speed of four contemporary ×86 simulators: gem5, Multi2sim, PTLsim and Sniper. We also discuss the strengths and limitations of these simulators. We believe that this paper provides insights into different simulation strategies and aims to help computer architects understand the differences among existing simulation tools. Ayaz Akram, Lina Sawalha |
ICCD | 1 |