Lavanya Subramanian

dblp:10/10980 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Memory systems · 52% Hardware accelerators and domain-specific architectures · 13% Cloud and datacenter computing · 11%
Artificial intelligence
1 paper
Robot navigation and mapping · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 74% Environmental and earth informatics · 26%
Human-computer interaction and pervasive computing
1 paper
Design research and methods · 100%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 81% Operating systems · 19%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Design research and methods
participatory design
1.012026
Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community Orchard · CHI 2026
Memory systems › memory controller
memory scheduling
0.952016
BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016
DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware Accelerators · ACM Trans. Archit. Code Optim. 2016
MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013
Robotics › Robot navigation and mapping
SLAM
0.812024
SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024
Robotics › Robot navigation and mapping › SLAM › multi-sensor SLAM
visual-inertial SLAM
0.812024
SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024
Memory systems
DRAM
0.632018
Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018
Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013
Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.412020
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020
Bioinformatics and computational biology › sequence analysis
read mapping
0.412020
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
genome sequence analysis accelerator
0.412020
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020
Cloud and datacenter computing
cluster resource management and scheduling
0.412019
GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019
Cloud and datacenter computing › multi-tenancy
multi-tenant scheduling
0.412019
GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019
Memory systems › memory access latency
DRAM latency
0.312018
Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018
Integrated circuit design › memory circuit design
sense amplifier design
0.312018
Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018
Environmental and earth informatics
environmental monitoring
0.312026
Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community Orchard · CHI 2026
Memory systems
memory controller
0.322012
Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012
Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011
Memory systems
memory interference
0.322012
Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012
Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011
Embedded and real-time systems › real-time scheduling
deadline-aware scheduling
0.212016
DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware Accelerators · ACM Trans. Archit. Code Optim. 2016
Embedded and real-time systems
runtime adaptation
0.212024
SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024
Hardware accelerators and domain-specific architectures › robotics accelerator
SLAM accelerator
0.212024
SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024
Memory systems
cache
0.212015
The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015
Memory systems › cache management
cache interference
0.212015
The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015
Performance modeling and evaluation
workload characterization
0.212015
The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015
Memory systems › DRAM
DRAM architecture
0.212013
Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013
Memory systems › DRAM › DRAM latency reduction
low-latency DRAM
0.212013
Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013
Performance modeling and evaluation › performance prediction
slowdown estimation
0.212013
MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
integrated CPU-GPU system
0.112012
Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012
Memory systems
inter-application interference
0.122016
BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016
The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015
Memory systems › memory bandwidth
memory bandwidth optimization
0.112020
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020
Memory systems › on-chip memory
on-chip SRAM
0.112020
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020
Processor architecture and microarchitecture
multicore design
0.122016
BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016
Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012
Services computing and microservices
microservice architecture
0.112019
GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019

Methods — techniques the papers use, named apart from their topics

pose estimation · 1.5adaptive runtime · 1.5participatory design workshops · 1.0participatory design workshop · 1.0systolic array · 0.9SLA-aware scheduling · 0.8bitvector algorithm · 0.4bit vector algorithm · 0.4simultaneous read and precharge · 0.3circuit simulation · 0.3worst-case memory access time estimation · 0.2workload characterization · 0.2RTL implementation · 0.2
YearPublicationVenuePosition
2026 Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community Orchard
abstract
Climate change is intensifying extreme heat, prompting cities to deploy environmental sensor networks to capture hyperlocal conditions. However, top-down deployments that produce data without public input often fail to align with community needs. This paper presents a participatory design study of a high-density environmental sensor network co-developed with a volunteer-run community orchard as part of a city-wide system for heat resilience. Through two participatory design workshops, orchard volunteers acted as co-designers by collaboratively defining the network’s purpose, selecting sensor locations, and identifying key environmental data outputs. The workshops functioned as sites of infrastructuring—building relationships, technical literacy, and shared understanding—while situating environmental data within the orchard’s place-based practices of stewardship. From this process, we derive design criteria for community-driven sensor networks that prioritize both technical function and public formation. These contributions extend participatory design approaches in HCI and offer guidance for future deployments of environmental sensing technologies in community and urban agriculture contexts.
Dana Habeeb, Lavanya Subramanian, Rahul Devajji, Nick Polak
CHI2
2024 SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping
abstract
Simultaneous localization and mapping (SLAM) algorithms track an agent's movements through an unknown environment. SLAM must be fast and accurate to avoid adverse effects such as motion sickness in AR/VR headsets and navigation errors in autonomous robots and drones. However, accurate SLAM is computationally expensive and target platforms are often highly constrained. Therefore, to maintain real-time functionality, designers must either pay a large up-front cost to design specialized accelerators or reduce the algorithm's functionality, resulting in poor pose estimation.
Armand Behroozi, Vlad Fruchter, Lavanya Subramanian, Sriseshan Srikanth, Scott A. Mahlke
ASPLOS (3)4
2023 FARSI: An Early-stage Design Space Exploration Framework to Tame the Domain-specific System-on-chip Complexity
abstract
Domain-specific SoCs (DSSoCs) are an attractive solution for domains with extremely stringent power, performance, and area constraints. However, DSSoCs suffer from two fundamental complexities. On the one hand, their many specialized hardware blocks result in complex systems and thus high development effort. On the other hand, their many system knobs expand the complexity of design space, making the search for the optimal design difficult. Thus to reach prevalence, taming such complexities is necessary. To address these challenges, in this work, we identify the necessary features of an early-stage design space exploration framework that targets the complex design space of DSSoCs and provide an instance of one such framework that we refer to as FARSI. FARSI provides an agile system-level simulator with speed up and accuracy of 8,400× and 98.5% compared to Synopsys Platform Architect. FARSI also provides an efficient exploration heuristic and achieves up to 62× and 35× improvement in convergence time compared to the classic simulated annealing (SA) and modern Multi-Objective Optimistic Search. This is done by augmenting SA with architectural reasoning such as locality exploitation and bottleneck relaxation. Furthermore, we embed various co-design capabilities and show that, on average, they have a 32% impact on the convergence rate. Finally, we demonstrate that using development-cost-aware policies can lower the system complexity, both in terms of the component count and variation by as much as 60% and 82% (e.g., for Network-on-a-Chip subsystem), respectively.
Behzad Boroujerdian, Devashree Tripathy, Lavanya Subramanian, Luke Yen, Vincent Lee, Vivek Venkatesan, Amit Jindal, Robert Shearer, Vijay Janapa Reddi
ACM Trans. Embed. Comput. Syst.5
2020 GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis
abstract
Genome sequence analysis has enabled significant advancements in medical and scientific areas such as personalized medicine, outbreak tracing, and the understanding of evolution. To perform genome sequencing, devices extract small random fragments of an organism's DNA sequence (known as reads). The first step of genome sequence analysis is a computational process known as read mapping. In read mapping, each fragment is matched to its potential location in the reference genome with the goal of identifying the original location of each read in the genome. Unfortunately, rapid genome sequencing is currently bottlenecked by the computational power and memory bandwidth limitations of existing systems, as many of the steps in genome sequence analysis must process a large amount of data. A major contributor to this bottleneck is approximate string matching (ASM), which is used at multiple points during the mapping process. ASM enables read mapping to account for sequencing errors and genetic variations in the reads. We propose GenASM, the first ASM acceleration framework for genome sequence analysis. GenASM performs bitvectorbased ASM, which can efficiently accelerate multiple steps of genome sequence analysis. We modify the underlying ASM algorithm (Bitap) to significantly increase its parallelism and reduce its memory footprint. Using this modified algorithm, we design the first hardware accelerator for Bitap. Our hardware accelerator consists of specialized systolic-array-based compute units and on-chip SRAMs that are designed to match the rate of computation with memory capacity and bandwidth, resulting in an efficient design whose performance scales linearly as we increase the number of compute units working in parallel. We demonstrate that GenASM provides significant performance and power benefits for three different use cases in genome sequence analysis. First, GenASM accelerates read alignment for both long reads and short reads. For long reads, GenASM outperforms state-of-the-art software and hardware accelerators by 116× and 3.9×, respectively, while reducing power consumption by 37× and 2.7×. For short reads, GenASM outperforms state-of-the-art software and hardware accelerators by 111× and 1.9×. Second, GenASM accelerates pre-alignment filtering for short reads, with 3.7× the performance of a state-of-the-art pre-alignment filter, while reducing power consumption by 1.7× and significantly improving the filtering accuracy. Third, GenASM accelerates edit distance calculation, with 22-12501× and 9.3-400× speedups over the state-of-the-art software library and FPGA-based accelerator, respectively, while reducing power consumption by 548-582× and 67×. We conclude that GenASM is a flexible, high-performance, and low-power framework, and we briefly discuss four other use cases that can benefit from GenASM.
Damla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina, Lavanya Subramanian, Jeremie S. Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan Gómez-Luna, Amirali Boroumand, Anant Nori, Allison Scibisz, Sreenivas Subramoney, Can Alkan, Saugata Ghose, Onur Mutlu
MICRO5
2019 GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks
abstract
The microservice architecture has dramatically reduced user effort in adopting and maintaining servers by providing a catalog of functions as services that can be used as building blocks to construct applications. This has enabled datacenter operators to look at managing datacenter hosting microservices quite differently from traditional infrastructures. Such a paradigm shift calls for a need to rethink resource management strategies employed in such execution environments. We observe that the visibility enabled by a microservices execution framework can be exploited to achieve high throughput and resource utilization while still meeting Service Level Agreements, especially in multi-tenant execution scenarios.
Ram Srivatsa Kannan, Lavanya Subramanian, Ashwin Raju, Jeongseob Ahn, Jason Mars, Lingjia Tang
EuroSys2
2018 Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems
abstract
DRAM memory access is a critical performance bottleneck. To access one cache block, an entire row needs to be sensed and amplified, data restored into the bitcells and the bitlines precharged, incurring high latency. Isolating the bitlines and sense amplifiers after activation enables reads and precharges to happen in parallel. However, there are challenges in achieving this isolation. We tackle these challenges and propose an effective scheme, simultaneous read and precharge (SRP), to isolate the sense amplifiers and bitlines and serve reads and precharges in parallel. Our detailed architecture and circuit simulations demonstrate that our simultaneous read and precharge (SRP) mechanism is able to achieve an 8.6% performance benefit over baseline, while reducing sense amplifier idle power by 30%, as compared to prior work, over a wide range of workloads.
Lavanya Subramanian, Kaushik Vaidyanathan, Anant Nori, Sreenivas Subramoney, Tanay Karnik, Hong Wang 0003
DAC1
2016 DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware Accelerators
abstract
Modern SoCs integrate multiple CPU cores and hardware accelerators (HWAs) that share the same main memory system, causing interference among memory requests from different agents. The result of this interference, if it is not controlled well, is missed deadlines for HWAs and low CPU performance. Few previous works have tackled this problem. State-of-the-art mechanisms designed for CPU-GPU systems strive to meet a target frame rate for GPUs by prioritizing the GPU close to the time when it has to complete a frame. We observe two major problems when such an approach is adapted to a heterogeneous CPU-HWA system. First, HWAs miss deadlines because they are prioritized only when close to their deadlines. Second, such an approach does not consider the diverse memory access characteristics of different applications running on CPUs and HWAs, leading to low performance for latency-sensitive CPU applications and deadline misses for some HWAs, including GPUs. In this article, we propose a Deadline-Aware memory Scheduler for Heterogeneous systems (DASH), which overcomes these problems using three key ideas, with the goal of meeting HWAs’ deadlines while providing high CPU performance. First, DASH prioritizes an HWA when it is not on track to meet its deadline any time during a deadline period, instead of prioritizing it only when close to a deadline. Second, DASH prioritizes HWAs over memory-intensive CPU applications based on the observation that memory-intensive applications’ performance is not sensitive to memory latency. Third, DASH treats short-deadline HWAs differently as they are more likely to miss their deadlines and schedules their requests based on worst-case memory access time estimates. Extensive evaluations across a wide variety of different workloads and systems show that DASH achieves significantly better CPU performance than the best previous scheduler while always meeting the deadlines for all HWAs, including GPUs, thereby largely improving frame rates.
Hiroyuki Usui, Lavanya Subramanian, Kevin Kai-Wei Chang, Onur Mutlu
ACM Trans. Archit. Code Optim.2
2016 BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling
abstract
In a multicore system, applications running on different cores interfere at main memory. This inter-application interference degrades overall system performance and unfairly slows down applications. Prior works have developed application-aware memory request schedulers to tackle this problem. State-of-the-art application-aware memory request schedulers prioritize memory requests of applications that are vulnerable to interference, by ranking individual applications based on their memory access characteristics and enforcing a total rank order. In this paper, we observe that state-of-the-art application-aware memory schedulers have two major shortcomings. First, such schedulers trade off hardware complexity in order to achieve high performance or fairness, since ranking applications individually with a total order based on memory access characteristics leads to high hardware cost and complexity. Such complexity could prevent the scheduler from meeting the stringent timing requirements of state-of-the-art DDR protocols. Second, ranking can unfairly slow down applications that are at the bottom of the ranking stack, thereby sometimes leading to high slowdowns and low overall system performance. To overcome these shortcomings, we propose the Blacklisting Memory Scheduler (BLISS), which achieves high system performance and fairness while incurring low hardware cost and complexity. BLISS design is based on two new observations. First, we find that, to mitigate interference, it is sufficient to separate applications into only two groups, one containing applications that are vulnerable to interference and another containing applications that cause interference, instead of ranking individual applications with a total order. Vulnerable-to-interference group is prioritized over the interference-causing group. Second, we show that this grouping can be efficiently performed by simply counting the number of consecutive requests served from each application. We evaluate BLISS across a wide variety of workloads and system configurations and compare its performance and hardware complexity (via RTL implementations), with five state-of-the-art memory schedulers. Our evaluations show that BLISS achieves 5 percent better system performance and 25 percent better fairness than the best-performing previous memory scheduler while greatly reducing critical path latency and hardware area cost of the memory scheduler (by 79 and 43 percent, respectively), thereby achieving a good trade-off between performance, fairness and hardware complexity.
Lavanya Subramanian, Donghyuk Lee, Vivek Seshadri, Harsha Rastogi, Onur Mutlu
IEEE Trans. Parallel Distributed Syst.1
2015 Decoupled Direct Memory Access: Isolating CPU and IO Traffic by Leveraging a Dual-Data-Port DRAM
abstract
Memory channel contention is a critical performance bottleneck in modern systems that have highly parallelized processing units operating on large data sets. The memory channel is contended not only by requests from different user applications (CPU access) but also by system requests for peripheral data (IO access), usually controlled by Direct Memory Access (DMA) engines. Our goal, in this work, is to improve system performance byeliminating memory channel contention between CPU accesses and IO accesses. To this end, we propose a hardware-software cooperative data transfer mechanism, Decoupled DMA (DDMA) that provides a specialized low-cost memory channel for IO accesses. In our DDMA design, main memoryhas two independent data channels, of which one is connected to the processor (CPU channel) and the other to the IO devices (IO channel), enabling CPU and IO accesses to be served on different channels. Systemsoftware or the compiler identifies which requests should be handled on the IO channel and communicates this to the DDMA engine, which then initiates the transfers on the IO channel. By doing so, our proposal increasesthe effective memory channel bandwidth, thereby either accelerating data transfers between system components, or providing opportunities to employ IO performance enhancement techniques (e.g., aggressive IO prefetching)without interfering with CPU accessesWe demonstrate the effectiveness of our DDMA framework in two scenarios: (i) CPU-GPU communication and (ii) in-memory communication (bulk datacopy/initialization within the main memory). By effectively decoupling accesses for CPU-GPU communication and in-memory communication from CPU accesses, our DDMA-based design achieves significant performanceimprovement across a wide variety of system configurations (e.g., 20% average performance improvement on a typical 2-channel 2-rank memory system).
Donghyuk Lee, Lavanya Subramanian, Rachata Ausavarungnirun, Jongmoo Choi, Onur Mutlu
PACT2
2015 The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory
abstract
In a multi-core system, interference at shared resources (such as caches and main memory) slows down applications running on different cores. Accurately estimating the slowdown of each application has several benefits: e.g., it can enable shared resource allocation in a manner that avoids unfair application slowdowns or provides slowdown guarantees. Unfortunately, prior works on estimating slowdowns either lead to inaccurate estimates, do not take into account shared caches, or rely on a priori application knowledge. This severely limits their applicability.
Lavanya Subramanian, Vivek Seshadri, Samira Manabi Khan, Onur Mutlu
MICRO1
2015 A-DRM: Architecture-aware Distributed Resource Management of Virtualized Clusters
abstract
Virtualization technologies has been widely adopted by large-scale cloud computing platforms. These virtualized systems employ distributed resource management (DRM) to achieve high resource utilization and energy savings by dynamically migrating and consolidating virtual machines. DRM schemes usually use operating-system-level metrics, such as CPU utilization, memory capacity demand and I/O utilization, to detect and balance resource contention. However, they are oblivious to microarchitecture-level resource interference (e.g., memory bandwidth contention between different VMs running on a host), which is currently not exposed to the operating system.
Canturk Isci, Lavanya Subramanian, Jongmoo Choi, Depei Qian 0001, Onur Mutlu
VEE3
2014 The Blacklisting Memory Scheduler: Achieving high performance and fairness at low cost
abstract
In a multicore system, applications running on different cores interfere at main memory. This inter-application interference degrades overall system performance and unfairly slows down applications. Prior works have developed application-aware memory request schedulers to tackle this problem. State-of-the-art application-aware memory request schedulers prioritize memory requests of applications that are vulnerable to interference, by ranking individual applications based on their memory access characteristics and enforcing a total rank order. In this paper, we observe that state-of-the-art application-aware memory schedulers have two major shortcomings. First, ranking applications individually with a total order based on memory access characteristics leads to high hardware cost and complexity. Second, ranking can unfairly slow down applications that are at the bottom of the ranking stack. To overcome these shortcomings, we propose the Blacklisting Memory Scheduler (BLISS), which achieves high system performance and fairness while incurring low hardware cost and complexity. BLISS design is based on two new observations. First, we find that, to mitigate interference, it is sufficient to separate applications into only two groups, one containing applications that cause interference and another containing applications vulnerable to interference, instead of ranking individual applications with a total order. Vulnerable-to-interference group is prioritized over the interference-causing group. Second, we show that this grouping can be efficiently performed by simply counting the number of consecutive requests served from each application - an application that has a large number of consecutive requests served is dynamically classified as interference-causing. We evaluate BLISS across a wide variety of workloads and system configurations and compare its performance and complexity with five state-of-the-art memory schedulers. Our evaluations show that BLISS achieves 5% better system performance and 25% better fairness than the best-performing previous memory scheduler while greatly reducing critical path latency and hardware area cost of the memory scheduler (by 79% and 43%, respectively).
Lavanya Subramanian, Donghyuk Lee, Vivek Seshadri, Harsha Rastogi, Onur Mutlu
ICCD1
2013 Tiered-latency DRAM: A low latency and low cost DRAM architecture
abstract
The capacity and cost-per-bit of DRAM have historically scaled to satisfy the needs of increasingly large and complex computer systems. However, DRAM latency has remained almost constant, making memory latency the performance bottleneck in today's systems. We observe that the high access latency is not intrinsic to DRAM, but a trade-off made to decrease cost-per-bit. To mitigate the high area overhead of DRAM sensing structures, commodity DRAMs connect many DRAM cells to each sense-amplifier through a wire called a bitline. These bitlines have a high parasitic capacitance due to their long length, and this bitline capacitance is the dominant source of DRAM latency. Specialized low-latency DRAMs use shorter bitlines with fewer cells, but have a higher cost-per-bit due to greater sense-amplifier area overhead. In this work, we introduce Tiered-Latency DRAM (TL-DRAM), which achieves both low latency and low cost-per-bit. In TL-DRAM, each long bitline is split into two shorter segments by an isolation transistor, allowing one segment to be accessed with the latency of a short-bitline DRAM without incurring high cost-per-bit. We propose mechanisms that use the low-latency segment as a hardware-managed or software-managed cache. Evaluations show that our proposed mechanisms improve both performance and energy-efficiency for both single-core and multi-programmed workloads.
Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian, Onur Mutlu
HPCA5
2013 MISE: Providing performance predictability and improving fairness in shared main memory systems
abstract
Applications running concurrently on a multicore system interfere with each other at the main memory. This interference can slow down different applications differently. Accurately estimating the slow down of each application in such a system can enable mechanisms that can enforce quality-of-service. While much prior work has focused on mitigating the performance degradation due to inter-application interference, there is little work on estimating slow down of individual applications in a multi-programmed environment. Our goal in this work is to build such an estimation scheme. To this end, we present our simple Memory-Interference-induced Slowdown Estimation (MISE) model that estimates slowdowns caused by memory interference. We build our model based on two observations. First, the performance of a memory-bound application is roughly proportional to the rate at which its memory requests are served, suggesting that request-service-rate can be used as a proxy for performance. Second, when an application's requests are prioritized over all other applications' requests, the application experiences very little interference from other applications. This provides a means for estimating the uninterfered request-service-rate of an application while it is run alongside other applications. Using the above observations, our model estimates the slowdown of an application as the ratio of its uninterfered and interfered request service rates. We propose simple changes to the above model to estimate the slowdown of non-memory-bound applications. We demonstrate the effectiveness of our model by developing two new memory scheduling schemes: 1) one that provides soft quality-of-service guarantees and 2) another that explicitly attempts to minimize maximum slowdown (i.e., unfairness) in the system. Evaluations show that our techniques perform significantly better than state-of-the-art memory scheduling approaches to address the above problems.
Lavanya Subramanian, Vivek Seshadri, Yoongu Kim, Ben Jaiyen, Onur Mutlu
HPCA1
2012 Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems
abstract
When multiple processor (CPU) cores and a GPU integrated together on the same chip share the off-chip main memory, requests from the GPU can heavily interfere with requests from the CPU cores, leading to low system performance and starvation of CPU cores. Unfortunately, state-of-the-art application-aware memory scheduling algorithms are ineffective at solving this problem at low complexity due to the large amount of GPU traffic. A large and costly request buffer is needed to provide these algorithms with enough visibility across the global request stream, requiring relatively complex hardware implementations. This paper proposes a fundamentally new approach that decouples the memory controller's three primary tasks into three significantly simpler structures that together improve system performance and fairness, especially in integrated CPU-GPU systems. Our three-stage memory controller first groups requests based on row-buffer locality. This grouping allows the second stage to focus only on inter-application request scheduling. These two stages enforce high-level policies regarding performance and fairness, and therefore the last stage consists of simple per-bank FIFO queues (no further command reordering within each bank) and straightforward logic that deals only with low-level DRAM commands and timing. We evaluate the design trade-offs involved in our Staged Memory Scheduler (SMS) and compare it against three state-of-the-art memory controller designs. Our evaluations show that SMS improves CPU performance without degrading GPU frame rate beyond a generally acceptable level, while being significantly less complex to implement than previous application-aware schedulers. Furthermore, SMS can be configured by the system software to prioritize the CPU or the GPU at varying levels to address different performance needs.
Rachata Ausavarungnirun, Kevin Kai-Wei Chang, Lavanya Subramanian, Gabriel H. Loh, Onur Mutlu
ISCA3
2011 Reducing memory interference in multicore systems via application-aware memory channel partitioning
abstract
Main memory is a major shared resource among cores in a multicore system. If the interference between different applications' memory requests is not controlled effectively, system performance can degrade significantly. Previous work aimed to mitigate the problem of interference between applications by changing the scheduling policy in the memory controller, i.e., by prioritizing memory requests from applications in a way that benefits system performance.
Sai Prashanth Muralidhara, Lavanya Subramanian, Onur Mutlu, Mahmut T. Kandemir, Thomas Moscibroda
MICRO2