Muhammad Laghari

dblp:332/2132 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0002-5661-6168ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 100%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
memory compression
2.132024
Memory Allocation Under Hardware Compression · MICRO 2024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Translation-optimized Memory Compression for Capacity · MICRO 2022
Memory systems › memory compression
hardware compressed memory
1.522024
Memory Allocation Under Hardware Compression · MICRO 2024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Memory systems › memory management › virtual memory
address translation
1.322024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Translation-optimized Memory Compression for Capacity · MICRO 2022
Operating systems › resource management › memory management
memory allocation
0.812024
Memory Allocation Under Hardware Compression · MICRO 2024
Operating systems › resource management
memory management
0.812024
Memory Allocation Under Hardware Compression · MICRO 2024
Memory systems › memory management › virtual memory › address translation
TLB
0.812024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Memory systems › memory management
virtual memory
0.812024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Memory systems
DRAM
0.632024
Memory Allocation Under Hardware Compression · MICRO 2024
DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory · ISCA 2024
Translation-optimized Memory Compression for Capacity · MICRO 2022
Memory systems › virtual memory management
page migration
0.212022
Translation-optimized Memory Compression for Capacity · MICRO 2022

Methods — techniques the papers use, named apart from their topics

MMU-like component design · 1.5FPGA prototype · 1.5design space exploration · 0.6deflate · 0.6
YearPublicationVenuePosition
2024 DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed Memory
abstract
To expand effective memory capacity, hardware memory compression transparently compresses and packs memory values more densely together in DRAM. This requires introducing a new layer of hardware-managed address translation in the memory controller (MC). However, for large and irregular workloads that already suffer from frequent virtual address translation misses in the TLB, adding an additional layer of address translation can double the translation misses (e.g., by adding a new miss in the MC per TLB miss). While TLB misses can be drastically reduced by using huge pages, no prior work has explored huge-page-like translation reach for hardware memory compression. While compressing and moving an entire huge page worth of data at a time can lead to huge-page-like address translation, moving a huge page worth of data together can consume an exorbitant amount of memory bandwidth.This paper explores how to achieve huge-page-like translation performance in this new address translation layer, while keeping compression at the page (instead of huge page) granularity. We propose dynamically shortening the translation entries of hot pages to only a few bits per entry by migrating hot pages to the limited number of DRAM locations whose addresses can be encoded using a few bits; colder pages still use the bigger fulllength translations so that colder pages can be placed anywhere in memory to fully utilize all the space in memory. Each short translation is tiny (e.g., 2 bits); as such, a 128KB translation cache filled mostly with short translations can achieve similar (e.g., 2GB) total translation reach as a TLB filled entirely with huge page entries. Evaluations show our idea – Dynamic Length Compressed-Memory Translations (DyLeCT) – improves average performance by 10.25% over the prior art.
Gagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun Jian 0002
ISCA2
2024 Memory Allocation Under Hardware Compression
abstract
As the scaling of memory density slows physically, a promising solution is to scale memory logically by enhancing the CPU's memory controller to encode and store data more densely in memory. This is known as hardware memory compression. Hardware memory compression decouples OS-managed physical memory from actual memory (i.e., DRAM); the memory controller spends a dynamically varying amount of DRAM on each physical page, depending on the compressibility of the page's content. The newly-decoupled actual memory effectively forms a new layer of memory beyond the traditional layers of virtual, pseudo-physical, and physical memory. We note unlike these traditional memory layers, each with its own specialized allocation interface (e.g., malloc/mmap for virtual memory, page tables+MMU for physical memory), this new layer of memory introduced by hardware memory compression still awaits its own unique memory allocation interface; its absence makes the allocation of actual memory imprecise and, sometimes, even impossible. Imprecisely allocating less actual memory, and/or unable to allocate more, can harm performance. Even imprecisely allocating more actual memory to some jobs can be harmful as it can result in allocating less actual memory to other jobs in highly-occupied memory systems, where compression is useful. To restore precise memory allocation, we design a new memory allocation specialized for this new layer of memory and, subsequently, architect a new MMU-like component in the memory controller and tackle the corresponding design challenges. We create a full-system FPGA prototype of a hardware-compressed memory system with precise memory allocation. Our evaluations using the prototype show that jobs perform stably under colocation. The performance variation is only 1%-2%; in comparison, it is 19%-89% under the prior art.
Muhammad Laghari, Gagandeep Panwar, David Bears, Chandler Jearls, Raghavendra Srinivas, Esha Choukse, Kirk W. Cameron, Ali Raza Butt, Xun Jian 0002
MICRO1
2022 Translation-optimized Memory Compression for Capacity
abstract
The demand for memory is ever increasing. Many prior works have explored hardware memory compression to increase effective memory capacity. However, prior works compress and pack/migrate data at a small - memory block-level - granularity; this introduces an additional block-level translation after the page-level virtual address translation. In general, the smaller the granularity of address translation, the higher the translation overhead. As such, this additional block-level translation exacerbates the well-known address translation problem for large and/or irregular workloads. A promising solution is to only save memory from cold (i.e., less recently accessed) pages without saving memory from hot (i.e., more recently accessed) pages (e.g., keep the hot pages uncompressed); this avoids block-level translation overhead for hot pages. However, it still faces two challenges. First, after a compressed cold page becomes hot again, migrating the page to a full 4KB DRAM location still adds another level (albeit page-level, instead of block-level) of translation on top of existing virtual address translation. Second, only compressing cold data require compressing them very aggressively to achieve high overall memory savings; decompressing very aggressively compressed data is very slow (e.g., $\gt 800 ns$ assuming the latest Deflate ASIC in industry). This paper presents Translation-optimized Memory Compression for Capacity (TMCC) to tackle the two challenges above. To address the first challenge, we propose compressing page table blocks in hardware to opportunistically embed compression translations into them in a software-transparent manner to effectively prefetch compression translations during a page walk, instead of serially fetching them after the walk. To address the second challenge, we perform a large design space exploration across many hardware configurations and diverse workloads to derive and implement in HDL an ASIC Deflate that is specialized for memory; for memory pages, it is 4X as fast as the state-of-the art ASIC Deflate, with little to no sacrifice in compression ratio. Our evaluations show that for large and/or irregular workloads, TMCC can either improve performance by 14% without sacrificing effective capacity or provide 2.2x the effective capacity without sacrificing performance compared to a state-of-the-art hardware memory compression for capacity.
Gagandeep Panwar, Muhammad Laghari, David Bears, Chandler Jearls, Esha Choukse, Kirk W. Cameron, Ali Raza Butt, Xun Jian 0002
MICRO2