Fazal Hameed

dblp:98/10051 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-2763-8755ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 12 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GNN and GRU-based dynamic NSGA-II algorithm to tolerate Byzantine faults for Software Defined Network
Nadir Shah, Gabriel-Miro Muntean, Waqar Mehmood, Fazal Hameed
J. Netw. Comput. Appl.5
2026 MoTO-BFT: Intelligent gradient-driven orchestration of BFT-aware controller placement and task offloading in vehicular edge SDN
abstract
Software-Defined Networking (SDN) has emerged as a key enabler for intelligent vehicular edge computing by providing centralized control and adaptive resource management. However, achieving reliable controller placement under Byzantine Fault Tolerance (BFT) and efficient task offloading in highly dynamic vehicular environments remains a major challenge due to frequent topology changes, communication delays, and control decision overhead. Existing works treat controller placement and task offloading as independent optimization problems and do not account for reassignment frequency, controller decision time, or BFT-compliant controller mapping. To address these limitations, this paper proposes MoTO-BFT , a novel M ulti- o bjective T ask O ffloading and B yzantine F ault- T olerant controller placement framework for SDN-enabled vehicular networks. MoTO-BFT integrates Temporal Convolutional Network (TCN)-based trajectory prediction with a novel Fast-converging Multi-objective Gradient Aggregation algorithm (FaMOGA) to jointly optimize BFT-compliant controller placement and computation task offloading. The proposed framework minimizes communication and computation delay, controller decision time, controller count, and reassignment frequency while ensuring balanced load distribution and 3 f + 1 BFT controller mapping. Extensive simulations using real vehicular mobility traces demonstrate that MoTO-BFT significantly outperforms existing controller placement and offloading baselines in terms of latency reduction, reassignment stability, and BFT robustness, making it a highly efficient and reliable solution for dynamic vehicular SDN environments.
Syed Aizaz Ul Haq, Nadir Shah, Fazal Hameed, Jan Badshah, Gabriel-Miro Muntean
J. Syst. Archit.4
2025 Scalability and Limitations of Existing Software Requirements Prioritization Techniques: A Systematic Literature Review
abstract
ABSTRACT Requirements prioritization puts more emphasis on the software requirements based on their importance, making it a crucial activity in managing software requirements throughout the software development process. This work explains the concept of prioritization techniques, their relevance, and related issues. Recent literature has described many requirement prioritization techniques, but each comes with its own limitations and challenges, specifically regarding scalability. In this article, we provide a systematic literature review (SLR) of the requirement prioritization techniques over the past decade and identify their limitations and scalability issues. In this context, we develop a planned protocol that includes all the necessary steps required for the SLR process. As a result of conducting the SLR, 53 primary studies were shortlisted for data extraction. The analysis of the results indicates that there is a minimal focus on large‐scale functional requirements, with only 9% of the work addressing this area. To the best of our knowledge, no work has been done on prioritizing requirements between ERP systems and mobile‐specific applications. Most existing studies focus on imaginary or simulated projects, and very limited actual work has been carried out during the execution of the industrial projects. When requirements are prioritized correctly, the likelihood of successful implementation will be high, and it will lead to a higher quality product.
Waqar Mehmood, Fazal Hameed, Muhammad Asif Nauman
J. Softw. Evol. Process.3
2023 An energy-efficient cache replacement policy for ultra-dense racetrack memory
Fazal Hameed, Moazam Maqsood, Syed Ali Irtaza
J. Syst. Archit.1
2023 ROLLED: Racetrack Memory Optimized Linear Layout and Efficient Decomposition of Decision Trees
abstract
Modern low power distributed systems tend to integrate machine learning algorithms. In resource-constrained setups, the execution of the models has to be optimized for performance and energy consumption. Racetrack memory (RTM) promises to achieve these goals by offering unprecedented integration density, smaller access latency, and reduced energy consumption. However, to access data in RTM, it needs to beshiftedto theaccess portfirst. We investigate decision trees and develop placement strategies to reduce the total number of shifts in RTM. Decision trees allow profiling during training, resulting in tree paths' access probabilities. We map tree nodes to RTM so that the total number of shifts is minimal. Concretely, we present two different placement approaches: 1) where tree nodes are closely packed and placeduniformlyin a single RTM location and 2) where decision tree nodes aredecomposedto separate RTM blocks. We discuss theoretical cost models for both approaches, we formally prove an upper bound of$4\times$for the unified and an upper bound of$12\times$for the decomposed organization towards the optimal placement. We conduct a thorough experimental evaluation to compare our algorithms to the state-of-the-art placement strategies Our experimental evaluations show that theunifiedanddecomposedsolutions reduce the number of shifts by$58.1\%$and$80.1\%$, respectively, leading to a$53.8\%$and$46.3\%$reduction in the overall runtime and$52.6\%$and$61.7\%$reduction in the energy consumption, compared to a naive baseline.
Christian Hakert, Asif Ali Khan, Kuan-Hsun Chen, Fazal Hameed, Jerónimo Castrillón, Jian-Jia Chen
IEEE Trans. Computers4
2023 DownShift: Tuning Shift Reduction With Reliability for Racetrack Memories
abstract
Ultra-dense non-volatileracetrack memories(RTMs) have been investigated at various levels in the memory hierarchy for improved performance and reduced energy consumption. However, the innateshiftoperations in RTMs, required for data access, incur performance penalties and can induce position errors. These factors can hinder their applicability in replacing low-latency, reliable on-chip memories. Intelligent placement of memory objects in RTMs can significantly reduce the number of shifts per memory access with little to no hardware overhead. However, existing placement strategies may lead to sub-optimal performance when applied to different architectures. Additionally, the impact of these shift optimization techniques on RTM reliability has been insufficiently investigated. We propose DownShift, a generalized data placement mechanism that improves upon prior approaches by taking into account (1) the timing and liveliness information of memory objects and (2) the underlying memory architecture, including required shifting fault tolerance. Thus, we also propose a collaboratively designed new shift alignment reliability technique called GROGU. GROGU leverages the reduced shift window made possible through DownShift allowing improved reliability, area, and energy compared to the state-of-the-art reliability approaches. DownShift reduces the number of shifts, runtime, and energy consumption by 3.24×, 47.6%, and 70.8% compared to the state-of-the-art. GROGU consumes 2.2× less area and 1.3× less energy while providing 16.8× improvement in shift fault tolerance compared to the leading reliability approach for a latency degradation of only 3.2%.
Asif Ali Khan, Sébastien Ollivier, Fazal Hameed, Jerónimo Castrillón, Alex K. Jones
IEEE Trans. Computers3
2022 BlendCache: An Energy and Area Efficient Racetrack Last-Level-Cache Architecture
abstract
Racetrack memory (RTM) is a promising nonvolatile memory that provides multibit storage cells achieving a higher area and leakage energy efficiency compared to contemporary volatile and nonvolatile memories. These features make RTM a potential candidate to be used as a last-level-cache (LLC). One drawback of the multibit RTM cell is the serialized access to the stored data, resulting in a shift penalty to access a particular bit within the cell. This overhead is particularly critical for LLC tags, for which prior RTM designs place tags either in SRAM or in single-bit RTM cells. While this avoids shifting, these designs require a large number of leaky cells incurring high energy consumption. To address this problem, this article proposes an energy-efficient RTM design called BlendCache that efficiently stores the tags in the leakage-optimized multibit RTM cells. To reduce the RTM shift penalty of these cells, BlendCache exploits the spatial locality of programs by maximizing accesses to nearby locations in RTM. Employing 32-bit RTM cells for a single core, BlendCache reduces the energy consumption by 20.8% and area by 15.2% compared to the state-of-the-art while its impact on performance is negligible. For a 4-core system, the energy improvement translates to 35.9% with 3% performance degradation.
Fazal Hameed, Jerónimo Castrillón
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 BLOwing Trees to the Ground: Layout Optimization of Decision Trees on Racetrack Memory
abstract
Modern distributed low power systems tend to integrate machine learning algorithms, which are directly executed on the distributed devices (on the edge). In resource constrained setups (e.g. battery driven sensor nodes), the execution of the machine learning models has to be optimized for execution time and energy consumption. Racetrack memory (RTM), an emerging non-volatile memory (NVM), promises to achieve these goals by offering unprecedented integration density, smaller access-latency and reduced energy consumption. However, in order to access data in RTM, it needs to be shifted to the access port first, resulting in latency and energy penalties. In this paper, we propose B.L.O. (Bidirectional Linear Ordering), a novel domain-specific approach for placing decision trees in RTMs. We reduce the total amount of shifts during inference by exploiting the tree structure and estimated access probabilities. We further apply the state-of-the-art methods to place data structures in RTM, without exploiting any domain-specific knowledge, to the decision trees and compare them to B. L.O. We formally prove that the B.L.O. solution has an approximation ratio of 4, i.e., its number of shifts is guaranteed to be at most 4 times the optimal number of shifts for a given decision tree. Throughout the experimental evaluation, we show that for the realistic use case B.L.O. empirically outperforms the state-of-the-art data placement method on average by 54.7% in terms of shifts, 19.2% in terms of runtime and 19.2% in terms of energy consumption.
Christian Hakert, Asif Ali Khan, Kuan-Hsun Chen, Fazal Hameed, Jerónimo Castrillón, Jian-Jia Chen
DAC4
2021 Improving the Performance of Block-based DRAM Caches Via Tag-Data Decoupling
abstract
In-package DRAM-based Last-Level-Caches (LLCs) that cache data in small chunks (i.e., blocks) are promising for improving system performance due to their efficient main memory bandwidth utilization. However, in these high-capacity DRAM caches, managing metadata (i.e., tags) at low cost is challenging. Storing the tags in SRAM has the advantage of quick tag access but is impractical due to a large area overhead. Storing the tags in DRAM reduces the area overhead but incurs tag serialization latency for an associative LLC design, which is inevitable for achieving high cache hit rate. To address the area and latency overhead problem, we propose a block-based DRAM LLC design that decouples tag and data into two regions in DRAM. Our design stores the tags in a latency-optimized DRAM region as the tags are accessed more often than the data. In contrast, we optimize the data region for area efficiency and map spatially-adjacent cache blocks to the same DRAM row to exploit spatial locality. Our design mitigates the tag serialization latency of existing associative DRAM LLCs via selective in-DRAM tag comparison, which overlaps the latency of tag and data accesses. This efficiently enables LLC bypassing via a novel DRAM Absence Table (DAT) that not only provides fast LLC miss detection but also reduces in-package bandwidth requirements. Our evaluation using SPEC2006 benchmarks shows that our tag-data decoupled LLC improves system performance by 11.7 percent compared to a state-of-the-art direct-mapped LLC design and by 7.2 percent compared to an existing associative LLC design.
Fazal Hameed, Asif Ali Khan, Jerónimo Castrillón
IEEE Trans. Computers1
2020 Generalized Data Placement Strategies for Racetrack Memories
abstract
Ultra-dense non-volatile racetrack memories (RTMs) have been investigated at various levels in the memory hierarchy for improved performance and reduced energy consumption. However, the innate shift operations in RTMs hinder their applicability to replace low-latency on-chip memories. Recent research has demonstrated that intelligent placement of memory objects in RTMs can significantly reduce the amount of shifts with no hardware overhead, albeit for specific system setups. However, existing placement strategies may lead to sub-optimal performance when applied to different architectures. In this paper we look at generalized data placement mechanisms that improve upon existing ones by taking into account the underlying memory architecture and the timing and liveliness information of memory objects. We propose a novel heuristic and a formulation using genetic algorithms that optimize key performance parameters. We show that, on average, our generalized approach improves the number of shifts, performance and energy consumption by 4.3 ×, 46% and 55% respectively compared to the state-of-the-art.
Asif Ali Khan, Andres Goens, Fazal Hameed, Jerónimo Castrillón
DATE3
2020 Magnetic Racetrack Memory: From Physics to the Cusp of Applications Within a Decade
abstract
Racetrack memory (RTM) is a novel spintronic memory-storage technology that has the potential to overcome fundamental constraints of existing memory and storage devices. It is unique in that its core differentiating feature is the movement of data, which is composed of magnetic domain walls (DWs), by short current pulses. This enables more data to be stored per unit area compared to any other current technologies. On the one hand, RTM has the potential for mass data storage with unlimited endurance using considerably less energy than today’s technologies. On the other hand, RTM promises an ultrafast nonvolatile memory competitive with static random access memory (SRAM) but with a much smaller footprint. During the last decade, the discovery of novel physical mechanisms to operate RTM has led to a major enhancement in the efficiency with which nanoscopic, chiral DWs can be manipulated. New materials and artificially atomically engineered thin-film structures have been found to increase the speed and lower the threshold current with which the data bits can be manipulated. With these recent developments, RTM has attracted the attention of the computer architecture community that has evaluated the use of RTM at various levels in the memory stack. Recent studies advocate RTM as a promising compromise between, on the one hand, power-hungry, volatile memories and, on the other hand, slow, nonvolatile storage. By optimizing the memory subsystem, significant performance improvements can be achieved, enabling a new era of cache, graphical processing units, and high capacity memory devices. In this article, we provide an overview of the major developments of RTM technology from both the physics and computer architecture perspectives over the past decade. We identify the remaining challenges and give an outlook on its future.
Robin Bläsing, Asif Ali Khan, Panagiotis Ch. Filippou, Chirag Garg, Fazal Hameed, Jerónimo Castrillón, Stuart S. P. Parkin
Proc. IEEE5
2020 ShiftsReduce: Minimizing Shifts in Racetrack Memory 4.0
abstract
Racetrack memories (RMs) have significantly evolved since their conception in 2008, making them a serious contender in the field of emerging memory technologies. Despite key technological advancements, the access latency and energy consumption of an RM-based system are still highly influenced by the number of shift operations. These operations are required to move bits to the right positions in the racetracks. This article presents data-placement techniques for RMs that maximize the likelihood that consecutive references access nearby memory locations at runtime, thereby minimizing the number of shifts. We present an integer linear programming (ILP) formulation for optimal data placement in RMs, and we revisit existing offset assignment heuristics, originally proposed for random-access memories. We introduce a novel heuristic tailored to a realistic RM and combine it with a genetic search to further improve the solution. We show a reduction in the number of shifts of up to 52.5%, outperforming the state of the art by up to 16.1%.
Asif Ali Khan, Fazal Hameed, Robin Bläsing, Stuart S. P. Parkin, Jerónimo Castrillón
ACM Trans. Archit. Code Optim.2
2020 Optimizing Tensor Contractions for Embedded Devices with Racetrack and DRAM Memories
abstract
Tensor contraction is a fundamental operation in many algorithms with a plethora of applications ranging from quantum chemistry over fluid dynamics and image processing to machine learning. The performance of tensor computations critically depends on the efficient utilization of on-chip/off-chip memories. In the context of low-power embedded devices, efficient management of the memory space becomes even more crucial, in order to meet energy constraints. This work aims at investigating strategies for performance- and energy-efficient tensor contractions on embedded systems, using racetrack memory (RTM)-based scratch-pad memory (SPM) and DRAM-based off-chip memory. Compiler optimizations such as the loop access order and data layout transformations paired with architectural optimizations such as prefetching and preshifting are employed to reduce the shifting overhead in RTMs. Optimizations for off-chip memory such as memory access order, data mapping and the choice of a suitable memory access granularity are employed to reduce the contention in the off-chip memory. Experimental results demonstrate that the proposed optimizations improve the SPM performance and energy consumption by 32% and 73%, respectively, compared to an iso-capacity SRAM. The overall DRAM dynamic energy consumption improvements due to memory optimizations amount to 80%.
Asif Ali Khan, Norman A. Rink, Fazal Hameed, Jerónimo Castrillón
ACM Trans. Embed. Comput. Syst.3
2019 SHRIMP: Efficient Instruction Delivery with Domain Wall Memory
abstract
Domain Wall Memory (DWM) is a promising emerging memory technology but suffers from the expensive shifts needed to align memory locations with access ports. Previous work on DWM concentrates on data, while, to the best of our knowledge, techniques to specifically target instruction streams have not yet been studied. In this paper, we propose Shift-Reducing Instruction Memory Placement (SHRIMP), the first instruction placement strategy suited for DWM which is accompanied with a supporting instruction fetch and memory architecture. The proposed approach reduces the number of shifts by 40% in the best case with a small memory overhead. In addition, SHRIMP achieves a best case of 23% reduction in total cycle counts.
Joonas Multanen, Pekka Jääskeläinen, Asif Ali Khan, Fazal Hameed, Jerónimo Castrillón
ISLPED4
2019 Optimizing tensor contractions for embedded devices with racetrack memory scratch-pads
abstract
Tensor contraction is a fundamental operation in many algorithms with a plethora of applications ranging from quantum chemistry over fluid dynamics and image processing to machine learning. The performance of tensor computations critically depends on the efficient utilization of on-chip memories. In the context of low-power embedded devices, efficient management of the memory space becomes even more crucial, in order to meet energy constraints. This work aims at investigating strategies for performance- and energy-efficient tensor contractions on embedded systems, using racetrack memory (RTM)-based scratch-pad memory (SPM). Compiler optimizations such as the loop access order and data layout transformations paired with architectural optimizations such as prefetching and preshifting are employed to reduce the shifting overhead in RTMs. Experimental results demonstrate that the proposed optimizations improve the SPM performance and energy consumption by 24% and 74% respectively compared to an iso-capacity SRAM.
Asif Ali Khan, Norman A. Rink, Fazal Hameed, Jerónimo Castrillón
LCTES3
2019 A Novel Hybrid DRAM/STT-RAM Last-Level-Cache Architecture for Performance, Energy, and Endurance Enhancement
abstract
High-capacity L4 architectures as a last-level cache (LLC) have been recently introduced between L3-SRAM and off-chip memory. These LLC architectures have either employed DRAM or spin-transfer torque (STT-RAM) memory technologies. It is a known fact that DRAM LLCs feature a higher energy consumption, while STT-RAM LLCs feature a lower write endurance compared to their counterparts. This paper proposes an efficient hybrid DRAM/STT-RAM LLC architecture that exploits the best characteristics offered by individual memory technologies while mitigating their drawbacks. More precisely, we introduce a novel mechanism for the storage and management of the hybrid LLC tags and a proactive L3-SRAM writeback policy that combines multiple dirty blocks that are mapped to the same LLC row. Our hybrid architecture reduces the LLC interference by having less writeback accesses and row fetches. The endurance is improved by reducing the number of STT-RAM block writes. We show that our LLC architecture reduces the total number of STT-RAM block writes by 78% and improves the average performance by 13% compared to a recently proposed STT-RAM LLC. Compared to the state-of-the-art DRAM LLC, we report an average energy and performance improvement of 24% and 17.1%, respectively.
Fazal Hameed, Jerónimo Castrillón
IEEE Trans. Very Large Scale Integr. Syst.1
2018 VAET-STT: Variation Aware STT-MRAM Analysis and Design Space Exploration Tool
abstract
Spin transfer torque magnetic random access memory is a promising candidate to replace CMOS based on-chip memories due to its advantages, such as nonvolatility, high density, and scalability. However, its stochastic switching and higher sensitivity to process variation compared to CMOS memories can significantly affect its performance, energy, and reliability. Although a few works exist which analyze the impact of process variation at the bit-cell level, such analysis at the system-level is missing. We have bridged this gap by developing a tool which can quantify the effect of stochasticity and process variations from the cell level to the overall memory system. The tool can perform a variation-aware design space exploration and memory configuration optimization for energy or performance while meeting reliability constraints. It also reports various failure rates and can evaluate the effectiveness of different error correcting code schemes. The results show that our framework can provide more realistic margins and the optimized variation-aware memory configuration could be significantly different from the conventional framework.
Sarath Mohanachandran Nair, Rajendra Bishnoi, Mohammad Saber Golanbari, Fabian Oboril, Fazal Hameed, Mehdi Baradaran Tahoori
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2018 Performance and Energy-Efficient Design of STT-RAM Last-Level Cache
Fazal Hameed, Asif Ali Khan, Jerónimo Castrillón
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Rethinking on-chip DRAM cache for simultaneous performance and energy optimization
abstract
State-of-the-art DRAM cache employs a small Tag-Cache and its performance is dependent upon two important parameters namely bank-level-parallelism and Tag-Cache hit rate. These parameters depend upon the row buffer organization. Recently, it has been shown that a small row buffer organization delivers better performance via improved bank-level-parallelism than the traditional large row buffer organization along with energy benefits. However, small row buffers do not fully exploit the temporal locality of tag accesses, leading to reduced TagCache hit rates. As a result, the DRAM cache needs to be re-designed for small row buffer organization to achieve additional performance benefits. In this paper, we propose a novel tag-store mechanism that improves the Tag-Cache hit rate by 70% compared to existing DRAM tag-store mechanisms employing small row buffer organization. In addition, we enhance the DRAM cache controller with novel policies that take into account the locality characteristics of cache accesses. We evaluate our novel tag-store mechanism and controller policies in an 8-core system running the SPEC2006 benchmark and compare their performance and energy consumption against recent proposals. Our architecture improves the average performance by 21.2% and 11.4% respectively compared to large and small row buffer organizations via simultaneously improving both parameters. Compared to DRAM cache with large row buffer organization, we report an energy improvement of 62%.
Fazal Hameed, Jerónimo Castrillón
DATE1
2016 Normally-OFF STT-MRAM Cache with Zero-Byte Compression for Energy Efficient Last-Level Caches
abstract
Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising alternative to SRAM due to its low leakage and scalability advantages. In fact, although being more energy-efficient than SRAM, STT-MRAM caches at higher levels (e.g. L3) still incur a high energy consumption due to 1) high leakage in their read and write circuits and 2) high dynamic write energy in their bit-cells. To address this problem, we propose a novel normally-off STT-MRAM cache that exploits the fact that most applications access zero-byte patterns very frequently. In this architecture, writing of zero-bytes is avoided to reduce write energy. In addition, all read and write circuits are by default power gated (i.e. normally-off) to reduce leakage power. Then, dynamically at runtime, only those circuits required for the ongoing operation are activated. Our evaluations for an L3-cache of a multi-core microprocessor show that this approach reduces the energy consumption by 60% compared to state-of-the-art, while its impact on performance is negligible.
Fabian Oboril, Fazal Hameed, Rajendra Bishnoi, Ali Ahari, Helia Naeimi, Mehdi Baradaran Tahoori
ISLPED2
2016 Architecting On-Chip DRAM Cache for Simultaneous Miss Rate and Latency Reduction
abstract
On-chip dynamic random access memory (DRAM) cache has been recently employed in the memory hierarchy to mitigate the widening latency gap between high-speed cores and off-chip memory. Two important parameters are the DRAM cache miss rate (D$-MR) and the DRAM cache hit latency (D$-HL), as they strongly influence the performance. These parameters depend upon the DRAM set mapping policy. Recently proposed DRAM set mapping policies are predominantly optimized for either D$-MR or D$-HL. We propose novel DRAM set mapping policies that simultaneously reduce D$-MR (via high associativity) and D$-HL (via improved row buffer hit rates). To further improve the D$-HL, we propose a small and low latency DRAM Tag cache (DTC) structure that can quickly determine whether an access to the DRAM cache will be a hit or a miss. The performance of the proposed DTC depends upon the DTC hit rate. To increase it, we present a novel DTC insertion policy that also increases the DTC hit rate. We investigate the latency and miss rate tradeoffs when designing a DRAM cache hierarchy and analyze the effects of different policies on the overall performance. We evaluate our policies on a wide variety of workloads and compare its performance with three recent proposals for on-chip DRAM caches. For a 16-core system, our set mapping policy along with our DTC and its adaptive DTC insertion policy improve the harmonic mean instruction per cycle throughput by 25.4%, 15.5%, and 7.3% compared to state-of-the-art, while requiring 55% less storage overhead for DRAM cache hit/miss prediction.
Fazal Hameed, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 Reducing Latency in an SRAM/DRAM Cache Hierarchy via a Novel Tag-Cache Architecture
abstract
Memory speed has become a major performance bottleneck as more and more cores are integrated on a multi-core chip. The widening latency gap between high speed cores and memory has led to the evolution of multi-level SRAM/DRAM cache hierarchies that exploit the latency benefits of smaller caches (e.g. private L1 and L2 SRAM caches) and the capacity benefits of larger caches (e.g. shared L3 SRAM and shared L4 DRAM cache). The main problem of employing large L3/L4 caches is their high tag lookup latency. To solve this problem, we introduce the novel concept of small and low latency SRAM/DRAM Tag-Cache structures that can quickly determine whether an access to the large L3/L4 caches will be a hit or a miss. The performance of the proposed Tag-Cache architecture depends upon the Tag-Cache hit rate and to improve it we propose a novel Tag-Cache insertion policy and a DRAM row buffer mapping policy that reduce the latency of memory requests. For a 16-core system, this improves the average harmonic mean instruction per cycle throughput of latency sensitive applications by 13.3% compared to state-of-the-art.
Fazal Hameed, Lars Bauer, Jörg Henkel
DAC1
2013 Simultaneously optimizing DRAM cache hit latency and miss rate via novel set mapping policies
abstract
Two key parameters that determine the performance of a DRAM cache based multi-core system are DRAM cache hit latency (HL) and DRAM cache miss rate (MR), as they strongly influence the average DRAM cache access latency. Recently proposed DRAM set mapping policies are either optimized for HL or for MR. None of these policies provides a good HL and MR at the same time. This paper presents a novel DRAM set mapping policy that simultaneously targets both parameters with the goal of achieving the best of both to reduce the overall DRAM cache access latency. For a 16-core system, our proposed set mapping policy reduces the average DRAM cache access latency (depends upon HL and MR) compared to state-of-the-art DRAM set mapping policies that are optimized for either HL or MR by 29.3% and 12.1%, respectively.
Fazal Hameed, Lars Bauer, Jörg Henkel
CASES1
2013 Adaptive cache management for a combined SRAM and DRAM cache hierarchy for multi-cores
abstract
On-chip DRAM caches may alleviate the memory bandwidth problem in future multi-core architectures through reducing off-chip accesses via increased cache capacity. For memory intensive applications, recent research has demonstrated the benefits of introducing high capacity on-chip L4-DRAM as Last-Level-Cache between L3-SRAM and off-chip memory. These multi-core cache hierarchies attempt to exploit the latency benefits of L3-SRAM and capacity benefits of L4-DRAM caches. However, not taking into consideration the cache access patterns of complex applications can cause inter-core DRAM interference and inter-core cache contention. In this paper, we contest to re-architect existing cache hierarchies by proposing a hybrid cache architecture, where the Last-Level-Cache is a combination of SRAM and DRAM caches. We propose an adaptive DRAM placement policy in response to the diverse requirements of complex applications with different cache access behaviors. It reduces inter-core DRAM interference and inter-core cache contention in SRAM/DRAM-based hybrid cache architectures: increasing the harmonic mean instruction-per-cycle throughput by 23.3% (max. 56%) and 13.3% (max. 35.1%) compared to state-of-the-art.
Fazal Hameed, Lars Bauer, Jörg Henkel
DATE1
2012 Dynamic cache management in multi-core architectures through run-time adaptation
abstract
Non-Uniform Cache Access (NUCA) architectures provide a potential solution to reduce the average latency for the last-level-cache (LLC), where the cache is organized into per-core local and remote partitions. Recent research has demonstrated the benefits of cooperative cache sharing among local and remote partitions. However, ignoring cache access patterns of concurrently executing applications sharing the local and remote partitions can cause inter-partition contention that reduces the overall instruction throughput. We propose a dynamic cache management scheme for LLC in NUCA-based architectures, which reduces inter-partition contention. Our proposed scheme provides efficient cache sharing by adapting migration, insertion, and promotion policies in response to the dynamic requirements of the individual applications with different cache access behaviors. Our adaptive cache management scheme allows individual cores to steal cache capacity from remote partitions to achieve better resource utilization. On average, our proposed scheme increases the performance (instructions per cycle) by 28% (minimum 8.4%, maximum 75%) compared to a private LLC organization.
Fazal Hameed, Lars Bauer, Jörg Henkel
DATE1
2011 Dynamic thermal management in 3D multi-core architecture through run-time adaptation
abstract
3D multi-core architectures are seen to provide increased transistor density, reduced power consumption, and improved performance through wire length reduction. However, 3D suffers from increased power density, which exacerbates thermal hotspots. In this paper, we present a novel 3D multi-core architecture that reduces processor activity on the die distant to the heat sink and a core-level dynamic thermal management technique based on the architectural adaptation, e.g. dynamically adapting core-resources depending on diverse application requirements and thermal behavior. The proposed thermal management technique synergistically combines the benefits of the architectural adaptation supported by our 3D multi-core architecture with dynamic voltage and frequency scaling. Our proposed technique provides 19.4% (maximum 24.4%, minimum 15.5%) improvement in the instruction throughput compared to the state-of-the-art thermal management techniques [4, 5] applied to the thermal-aware 3D processor architecture without considering run-time adaptation [10].
Fazal Hameed, Mohammad Abdullah Al Faruque, Jörg Henkel
DATE1