EDBT 2026 Demo / reviewers in the wild / expert
Lavanya Subramanian
dblp:10/10980
· DBLP profile ↗
16ranked-venue papers
5as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Memory systems · 52% Hardware accelerators and domain-specific architectures · 13% Cloud and datacenter computing · 11% | |
| Artificial intelligence
1 paper |
Robot navigation and mapping · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 74% Environmental and earth informatics · 26% | |
| Human-computer interaction and pervasive computing
1 paper |
Design research and methods · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Services computing and microservices · 81% Operating systems · 19% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Design research and methods
participatory design |
1.0 | 1 | 2026 | Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community Orchard · CHI 2026 |
Memory systems › memory controller
memory scheduling |
0.9 | 5 | 2016 | BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016 DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware Accelerators · ACM Trans. Archit. Code Optim. 2016 MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013 |
Robotics › Robot navigation and mapping
SLAM |
0.8 | 1 | 2024 | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024 |
Robotics › Robot navigation and mapping › SLAM › multi-sensor SLAM
visual-inertial SLAM |
0.8 | 1 | 2024 | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024 |
Memory systems
DRAM |
0.6 | 3 | 2018 | Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018 Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis |
0.4 | 1 | 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020 |
Bioinformatics and computational biology › sequence analysis
read mapping |
0.4 | 1 | 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020 |
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
genome sequence analysis accelerator |
0.4 | 1 | 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.4 | 1 | 2019 | GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019 |
Cloud and datacenter computing › multi-tenancy
multi-tenant scheduling |
0.4 | 1 | 2019 | GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019 |
Memory systems › memory access latency
DRAM latency |
0.3 | 1 | 2018 | Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018 |
Integrated circuit design › memory circuit design
sense amplifier design |
0.3 | 1 | 2018 | Closed yet open DRAM: achieving low latency and high performance in DRAM memory systems · DAC 2018 |
Environmental and earth informatics
environmental monitoring |
0.3 | 1 | 2026 | Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community Orchard · CHI 2026 |
Memory systems
memory controller |
0.3 | 2 | 2012 | Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012 Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Memory systems
memory interference |
0.3 | 2 | 2012 | Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012 Reducing memory interference in multicore systems via application-aware memory channel partitioning · MICRO 2011 |
Embedded and real-time systems › real-time scheduling
deadline-aware scheduling |
0.2 | 1 | 2016 | DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware Accelerators · ACM Trans. Archit. Code Optim. 2016 |
Embedded and real-time systems
runtime adaptation |
0.2 | 1 | 2024 | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024 |
Hardware accelerators and domain-specific architectures › robotics accelerator
SLAM accelerator |
0.2 | 1 | 2024 | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping · ASPLOS (3) 2024 |
Memory systems
cache |
0.2 | 1 | 2015 | The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015 |
Memory systems › cache management
cache interference |
0.2 | 1 | 2015 | The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2015 | The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015 |
Memory systems › DRAM
DRAM architecture |
0.2 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Memory systems › DRAM › DRAM latency reduction
low-latency DRAM |
0.2 | 1 | 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architecture · HPCA 2013 |
Performance modeling and evaluation › performance prediction
slowdown estimation |
0.2 | 1 | 2013 | MISE: Providing performance predictability and improving fairness in shared main memory systems · HPCA 2013 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
integrated CPU-GPU system |
0.1 | 1 | 2012 | Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012 |
Memory systems
inter-application interference |
0.1 | 2 | 2016 | BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016 The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memory · MICRO 2015 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.1 | 1 | 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020 |
Memory systems › on-chip memory
on-chip SRAM |
0.1 | 1 | 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis · MICRO 2020 |
Processor architecture and microarchitecture
multicore design |
0.1 | 2 | 2016 | BLISS: Balancing Performance, Fairness and Complexity in Memory Access Scheduling · IEEE Trans. Parallel Distributed Syst. 2016 Staged memory scheduling: Achieving high performance and scalability in heterogeneous systems · ISCA 2012 |
Services computing and microservices
microservice architecture |
0.1 | 1 | 2019 | GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks · EuroSys 2019 |
Methods — techniques the papers use, named apart from their topics
pose estimation · 1.5adaptive runtime · 1.5participatory design workshops · 1.0participatory design workshop · 1.0systolic array · 0.9SLA-aware scheduling · 0.8bitvector algorithm · 0.4bit vector algorithm · 0.4simultaneous read and precharge · 0.3circuit simulation · 0.3worst-case memory access time estimation · 0.2workload characterization · 0.2RTL implementation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Publics, Place, and Sensors: Co-Designing Environmental Monitoring with a Community OrchardabstractClimate change is intensifying extreme heat, prompting cities to deploy environmental sensor networks to capture hyperlocal conditions. However, top-down deployments that produce data without public input often fail to align with community needs. This paper presents a participatory design study of a high-density environmental sensor network co-developed with a volunteer-run community orchard as part of a city-wide system for heat resilience. Through two participatory design workshops, orchard volunteers acted as co-designers by collaboratively defining the network’s purpose, selecting sensor locations, and identifying key environmental data outputs. The workshops functioned as sites of infrastructuring—building relationships, technical literacy, and shared understanding—while situating environmental data within the orchard’s place-based practices of stewardship. From this process, we derive design criteria for community-driven sensor networks that prioritize both technical function and public formation. These contributions extend participatory design approaches in HCI and offer guidance for future deployments of environmental sensing technologies in community and urban agriculture contexts. Dana Habeeb, Lavanya Subramanian, Rahul Devajji, Nick Polak |
CHI | 2 |
| 2024 | SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and MappingabstractSimultaneous localization and mapping (SLAM) algorithms track an agent's movements through an unknown environment. SLAM must be fast and accurate to avoid adverse effects such as motion sickness in AR/VR headsets and navigation errors in autonomous robots and drones. However, accurate SLAM is computationally expensive and target platforms are often highly constrained. Therefore, to maintain real-time functionality, designers must either pay a large up-front cost to design specialized accelerators or reduce the algorithm's functionality, resulting in poor pose estimation. Armand Behroozi, Vlad Fruchter, Lavanya Subramanian, Sriseshan Srikanth, Scott A. Mahlke |
ASPLOS (3) | 4 |
| 2023 | FARSI: An Early-stage Design Space Exploration Framework to Tame the Domain-specific System-on-chip ComplexityabstractDomain-specific SoCs (DSSoCs) are an attractive solution for domains with extremely stringent power, performance, and area constraints. However, DSSoCs suffer from two fundamental complexities. On the one hand, their many specialized hardware blocks result in complex systems and thus high development effort. On the other hand, their many system knobs expand the complexity of design space, making the search for the optimal design difficult. Thus to reach prevalence, taming such complexities is necessary. To address these challenges, in this work, we identify the necessary features of an early-stage design space exploration framework that targets the complex design space of DSSoCs and provide an instance of one such framework that we refer to as FARSI. FARSI provides an agile system-level simulator with speed up and accuracy of 8,400× and 98.5% compared to Synopsys Platform Architect. FARSI also provides an efficient exploration heuristic and achieves up to 62× and 35× improvement in convergence time compared to the classic simulated annealing (SA) and modern Multi-Objective Optimistic Search. This is done by augmenting SA with architectural reasoning such as locality exploitation and bottleneck relaxation. Furthermore, we embed various co-design capabilities and show that, on average, they have a 32% impact on the convergence rate. Finally, we demonstrate that using development-cost-aware policies can lower the system complexity, both in terms of the component count and variation by as much as 60% and 82% (e.g., for Network-on-a-Chip subsystem), respectively. Behzad Boroujerdian, Devashree Tripathy, Lavanya Subramanian, Luke Yen, Vincent Lee, Vivek Venkatesan, Amit Jindal, Robert Shearer, Vijay Janapa Reddi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2020 | GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisabstractGenome sequence analysis has enabled significant advancements in medical and scientific areas such as personalized medicine, outbreak tracing, and the understanding of evolution. To perform genome sequencing, devices extract small random fragments of an organism's DNA sequence (known as reads). The first step of genome sequence analysis is a computational process known as read mapping. In read mapping, each fragment is matched to its potential location in the reference genome with the goal of identifying the original location of each read in the genome. Unfortunately, rapid genome sequencing is currently bottlenecked by the computational power and memory bandwidth limitations of existing systems, as many of the steps in genome sequence analysis must process a large amount of data. A major contributor to this bottleneck is approximate string matching (ASM), which is used at multiple points during the mapping process. ASM enables read mapping to account for sequencing errors and genetic variations in the reads. We propose GenASM, the first ASM acceleration framework for genome sequence analysis. GenASM performs bitvectorbased ASM, which can efficiently accelerate multiple steps of genome sequence analysis. We modify the underlying ASM algorithm (Bitap) to significantly increase its parallelism and reduce its memory footprint. Using this modified algorithm, we design the first hardware accelerator for Bitap. Our hardware accelerator consists of specialized systolic-array-based compute units and on-chip SRAMs that are designed to match the rate of computation with memory capacity and bandwidth, resulting in an efficient design whose performance scales linearly as we increase the number of compute units working in parallel. We demonstrate that GenASM provides significant performance and power benefits for three different use cases in genome sequence analysis. First, GenASM accelerates read alignment for both long reads and short reads. For long reads, GenASM outperforms state-of-the-art software and hardware accelerators by 116× and 3.9×, respectively, while reducing power consumption by 37× and 2.7×. For short reads, GenASM outperforms state-of-the-art software and hardware accelerators by 111× and 1.9×. Second, GenASM accelerates pre-alignment filtering for short reads, with 3.7× the performance of a state-of-the-art pre-alignment filter, while reducing power consumption by 1.7× and significantly improving the filtering accuracy. Third, GenASM accelerates edit distance calculation, with 22-12501× and 9.3-400× speedups over the state-of-the-art software library and FPGA-based accelerator, respectively, while reducing power consumption by 548-582× and 67×. We conclude that GenASM is a flexible, high-performance, and low-power framework, and we briefly discuss four other use cases that can benefit from GenASM. Damla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina, Lavanya Subramanian, Jeremie S. Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan Gómez-Luna, Amirali Boroumand, Anant Nori, Allison Scibisz, Sreenivas Subramoney, Can Alkan, Saugata Ghose, Onur Mutlu |
MICRO | 5 |
| 2019 | GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution FrameworksabstractThe microservice architecture has dramatically reduced user effort in adopting and maintaining servers by providing a catalog of functions as services that can be used as building blocks to construct applications. This has enabled datacenter operators to look at managing datacenter hosting microservices quite differently from traditional infrastructures. Such a paradigm shift calls for a need to rethink resource management strategies employed in such execution environments. We observe that the visibility enabled by a microservices execution framework can be exploited to achieve high throughput and resource utilization while still meeting Service Level Agreements, especially in multi-tenant execution scenarios. Ram Srivatsa Kannan, Lavanya Subramanian, Ashwin Raju, Jeongseob Ahn, Jason Mars, Lingjia Tang |
EuroSys | 2 |
| 2018 | Closed yet open DRAM: achieving low latency and high performance in DRAM memory systemsabstractDRAM memory access is a critical performance bottleneck. To access one cache block, an entire row needs to be sensed and amplified, data restored into the bitcells and the bitlines precharged, incurring high latency. Isolating the bitlines and sense amplifiers after activation enables reads and precharges to happen in parallel. However, there are challenges in achieving this isolation. We tackle these challenges and propose an effective scheme, simultaneous read and precharge (SRP), to isolate the sense amplifiers and bitlines and serve reads and precharges in parallel. Our detailed architecture and circuit simulations demonstrate that our simultaneous read and precharge (SRP) mechanism is able to achieve an 8.6% performance benefit over baseline, while reducing sense amplifier idle power by 30%, as compared to prior work, over a wide range of workloads. Lavanya Subramanian, Kaushik Vaidyanathan, Anant Nori, Sreenivas Subramoney, Tanay Karnik, Hong Wang 0003 |
DAC | 1 |
| 2016 | DASH: Deadline-Aware High-Performance Memory Scheduler for Heterogeneous Systems with Hardware AcceleratorsabstractModern SoCs integrate multiple CPU cores and hardware accelerators (HWAs) that share the same main memory system, causing interference among memory requests from different agents. The result of this interference, if it is not controlled well, is missed deadlines for HWAs and low CPU performance. Few previous works have tackled this problem. State-of-the-art mechanisms designed for CPU-GPU systems strive to meet a target frame rate for GPUs by prioritizing the GPU close to the time when it has to complete a frame. We observe two major problems when such an approach is adapted to a heterogeneous CPU-HWA system. First, HWAs miss deadlines because they are prioritized only when close to their deadlines. Second, such an approach does not consider the diverse memory access characteristics of different applications running on CPUs and HWAs, leading to low performance for latency-sensitive CPU applications and deadline misses for some HWAs, including GPUs. In this article, we propose a Deadline-Aware memory Scheduler for Heterogeneous systems (DASH), which overcomes these problems using three key ideas, with the goal of meeting HWAs’ deadlines while providing high CPU performance. First, DASH prioritizes an HWA when it is not on track to meet its deadline any time during a deadline period, instead of prioritizing it only when close to a deadline. Second, DASH prioritizes HWAs over memory-intensive CPU applications based on the observation that memory-intensive applications’ performance is not sensitive to memory latency. Third, DASH treats short-deadline HWAs differently as they are more likely to miss their deadlines and schedules their requests based on worst-case memory access time estimates. Extensive evaluations across a wide variety of different workloads and systems show that DASH achieves significantly better CPU performance than the best previous scheduler while always meeting the deadlines for all HWAs, including GPUs, thereby largely improving frame rates. Hiroyuki Usui, Lavanya Subramanian, Kevin Kai-Wei Chang, Onur Mutlu |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | BLISS: Balancing Performance, Fairness and Complexity in Memory Access SchedulingabstractIn a multicore system, applications running on different cores interfere at main memory. This inter-application interference degrades overall system performance and unfairly slows down applications. Prior works have developed application-aware memory request schedulers to tackle this problem. State-of-the-art application-aware memory request schedulers prioritize memory requests of applications that are vulnerable to interference, by ranking individual applications based on their memory access characteristics and enforcing a total rank order. In this paper, we observe that state-of-the-art application-aware memory schedulers have two major shortcomings. First, such schedulers trade off hardware complexity in order to achieve high performance or fairness, since ranking applications individually with a total order based on memory access characteristics leads to high hardware cost and complexity. Such complexity could prevent the scheduler from meeting the stringent timing requirements of state-of-the-art DDR protocols. Second, ranking can unfairly slow down applications that are at the bottom of the ranking stack, thereby sometimes leading to high slowdowns and low overall system performance. To overcome these shortcomings, we propose the Blacklisting Memory Scheduler (BLISS), which achieves high system performance and fairness while incurring low hardware cost and complexity. BLISS design is based on two new observations. First, we find that, to mitigate interference, it is sufficient to separate applications into only two groups, one containing applications that are vulnerable to interference and another containing applications that cause interference, instead of ranking individual applications with a total order. Vulnerable-to-interference group is prioritized over the interference-causing group. Second, we show that this grouping can be efficiently performed by simply counting the number of consecutive requests served from each application. We evaluate BLISS across a wide variety of workloads and system configurations and compare its performance and hardware complexity (via RTL implementations), with five state-of-the-art memory schedulers. Our evaluations show that BLISS achieves 5 percent better system performance and 25 percent better fairness than the best-performing previous memory scheduler while greatly reducing critical path latency and hardware area cost of the memory scheduler (by 79 and 43 percent, respectively), thereby achieving a good trade-off between performance, fairness and hardware complexity. Lavanya Subramanian, Donghyuk Lee, Vivek Seshadri, Harsha Rastogi, Onur Mutlu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Decoupled Direct Memory Access: Isolating CPU and IO Traffic by Leveraging a Dual-Data-Port DRAMabstractMemory channel contention is a critical performance bottleneck in modern systems that have highly parallelized processing units operating on large data sets. The memory channel is contended not only by requests from different user applications (CPU access) but also by system requests for peripheral data (IO access), usually controlled by Direct Memory Access (DMA) engines. Our goal, in this work, is to improve system performance byeliminating memory channel contention between CPU accesses and IO accesses. To this end, we propose a hardware-software cooperative data transfer mechanism, Decoupled DMA (DDMA) that provides a specialized low-cost memory channel for IO accesses. In our DDMA design, main memoryhas two independent data channels, of which one is connected to the processor (CPU channel) and the other to the IO devices (IO channel), enabling CPU and IO accesses to be served on different channels. Systemsoftware or the compiler identifies which requests should be handled on the IO channel and communicates this to the DDMA engine, which then initiates the transfers on the IO channel. By doing so, our proposal increasesthe effective memory channel bandwidth, thereby either accelerating data transfers between system components, or providing opportunities to employ IO performance enhancement techniques (e.g., aggressive IO prefetching)without interfering with CPU accessesWe demonstrate the effectiveness of our DDMA framework in two scenarios: (i) CPU-GPU communication and (ii) in-memory communication (bulk datacopy/initialization within the main memory). By effectively decoupling accesses for CPU-GPU communication and in-memory communication from CPU accesses, our DDMA-based design achieves significant performanceimprovement across a wide variety of system configurations (e.g., 20% average performance improvement on a typical 2-channel 2-rank memory system). Donghyuk Lee, Lavanya Subramanian, Rachata Ausavarungnirun, Jongmoo Choi, Onur Mutlu |
PACT | 2 |
| 2015 | The application slowdown model: quantifying and controlling the impact of inter-application interference at shared caches and main memoryabstractIn a multi-core system, interference at shared resources (such as caches and main memory) slows down applications running on different cores. Accurately estimating the slowdown of each application has several benefits: e.g., it can enable shared resource allocation in a manner that avoids unfair application slowdowns or provides slowdown guarantees. Unfortunately, prior works on estimating slowdowns either lead to inaccurate estimates, do not take into account shared caches, or rely on a priori application knowledge. This severely limits their applicability. Lavanya Subramanian, Vivek Seshadri, Samira Manabi Khan, Onur Mutlu |
MICRO | 1 |
| 2015 | A-DRM: Architecture-aware Distributed Resource Management of Virtualized ClustersabstractVirtualization technologies has been widely adopted by large-scale cloud computing platforms. These virtualized systems employ distributed resource management (DRM) to achieve high resource utilization and energy savings by dynamically migrating and consolidating virtual machines. DRM schemes usually use operating-system-level metrics, such as CPU utilization, memory capacity demand and I/O utilization, to detect and balance resource contention. However, they are oblivious to microarchitecture-level resource interference (e.g., memory bandwidth contention between different VMs running on a host), which is currently not exposed to the operating system. Canturk Isci, Lavanya Subramanian, Jongmoo Choi, Depei Qian 0001, Onur Mutlu |
VEE | 3 |
| 2014 | The Blacklisting Memory Scheduler: Achieving high performance and fairness at low costabstractIn a multicore system, applications running on different cores interfere at main memory. This inter-application interference degrades overall system performance and unfairly slows down applications. Prior works have developed application-aware memory request schedulers to tackle this problem. State-of-the-art application-aware memory request schedulers prioritize memory requests of applications that are vulnerable to interference, by ranking individual applications based on their memory access characteristics and enforcing a total rank order. In this paper, we observe that state-of-the-art application-aware memory schedulers have two major shortcomings. First, ranking applications individually with a total order based on memory access characteristics leads to high hardware cost and complexity. Second, ranking can unfairly slow down applications that are at the bottom of the ranking stack. To overcome these shortcomings, we propose the Blacklisting Memory Scheduler (BLISS), which achieves high system performance and fairness while incurring low hardware cost and complexity. BLISS design is based on two new observations. First, we find that, to mitigate interference, it is sufficient to separate applications into only two groups, one containing applications that cause interference and another containing applications vulnerable to interference, instead of ranking individual applications with a total order. Vulnerable-to-interference group is prioritized over the interference-causing group. Second, we show that this grouping can be efficiently performed by simply counting the number of consecutive requests served from each application - an application that has a large number of consecutive requests served is dynamically classified as interference-causing. We evaluate BLISS across a wide variety of workloads and system configurations and compare its performance and complexity with five state-of-the-art memory schedulers. Our evaluations show that BLISS achieves 5% better system performance and 25% better fairness than the best-performing previous memory scheduler while greatly reducing critical path latency and hardware area cost of the memory scheduler (by 79% and 43%, respectively). Lavanya Subramanian, Donghyuk Lee, Vivek Seshadri, Harsha Rastogi, Onur Mutlu |
ICCD | 1 |
| 2013 | Tiered-latency DRAM: A low latency and low cost DRAM architectureabstractThe capacity and cost-per-bit of DRAM have historically scaled to satisfy the needs of increasingly large and complex computer systems. However, DRAM latency has remained almost constant, making memory latency the performance bottleneck in today's systems. We observe that the high access latency is not intrinsic to DRAM, but a trade-off made to decrease cost-per-bit. To mitigate the high area overhead of DRAM sensing structures, commodity DRAMs connect many DRAM cells to each sense-amplifier through a wire called a bitline. These bitlines have a high parasitic capacitance due to their long length, and this bitline capacitance is the dominant source of DRAM latency. Specialized low-latency DRAMs use shorter bitlines with fewer cells, but have a higher cost-per-bit due to greater sense-amplifier area overhead. In this work, we introduce Tiered-Latency DRAM (TL-DRAM), which achieves both low latency and low cost-per-bit. In TL-DRAM, each long bitline is split into two shorter segments by an isolation transistor, allowing one segment to be accessed with the latency of a short-bitline DRAM without incurring high cost-per-bit. We propose mechanisms that use the low-latency segment as a hardware-managed or software-managed cache. Evaluations show that our proposed mechanisms improve both performance and energy-efficiency for both single-core and multi-programmed workloads. Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian, Onur Mutlu |
HPCA | 5 |
| 2013 | MISE: Providing performance predictability and improving fairness in shared main memory systemsabstractApplications running concurrently on a multicore system interfere with each other at the main memory. This interference can slow down different applications differently. Accurately estimating the slow down of each application in such a system can enable mechanisms that can enforce quality-of-service. While much prior work has focused on mitigating the performance degradation due to inter-application interference, there is little work on estimating slow down of individual applications in a multi-programmed environment. Our goal in this work is to build such an estimation scheme. To this end, we present our simple Memory-Interference-induced Slowdown Estimation (MISE) model that estimates slowdowns caused by memory interference. We build our model based on two observations. First, the performance of a memory-bound application is roughly proportional to the rate at which its memory requests are served, suggesting that request-service-rate can be used as a proxy for performance. Second, when an application's requests are prioritized over all other applications' requests, the application experiences very little interference from other applications. This provides a means for estimating the uninterfered request-service-rate of an application while it is run alongside other applications. Using the above observations, our model estimates the slowdown of an application as the ratio of its uninterfered and interfered request service rates. We propose simple changes to the above model to estimate the slowdown of non-memory-bound applications. We demonstrate the effectiveness of our model by developing two new memory scheduling schemes: 1) one that provides soft quality-of-service guarantees and 2) another that explicitly attempts to minimize maximum slowdown (i.e., unfairness) in the system. Evaluations show that our techniques perform significantly better than state-of-the-art memory scheduling approaches to address the above problems. Lavanya Subramanian, Vivek Seshadri, Yoongu Kim, Ben Jaiyen, Onur Mutlu |
HPCA | 1 |
| 2012 | Staged memory scheduling: Achieving high performance and scalability in heterogeneous systemsabstractWhen multiple processor (CPU) cores and a GPU integrated together on the same chip share the off-chip main memory, requests from the GPU can heavily interfere with requests from the CPU cores, leading to low system performance and starvation of CPU cores. Unfortunately, state-of-the-art application-aware memory scheduling algorithms are ineffective at solving this problem at low complexity due to the large amount of GPU traffic. A large and costly request buffer is needed to provide these algorithms with enough visibility across the global request stream, requiring relatively complex hardware implementations. This paper proposes a fundamentally new approach that decouples the memory controller's three primary tasks into three significantly simpler structures that together improve system performance and fairness, especially in integrated CPU-GPU systems. Our three-stage memory controller first groups requests based on row-buffer locality. This grouping allows the second stage to focus only on inter-application request scheduling. These two stages enforce high-level policies regarding performance and fairness, and therefore the last stage consists of simple per-bank FIFO queues (no further command reordering within each bank) and straightforward logic that deals only with low-level DRAM commands and timing. We evaluate the design trade-offs involved in our Staged Memory Scheduler (SMS) and compare it against three state-of-the-art memory controller designs. Our evaluations show that SMS improves CPU performance without degrading GPU frame rate beyond a generally acceptable level, while being significantly less complex to implement than previous application-aware schedulers. Furthermore, SMS can be configured by the system software to prioritize the CPU or the GPU at varying levels to address different performance needs. Rachata Ausavarungnirun, Kevin Kai-Wei Chang, Lavanya Subramanian, Gabriel H. Loh, Onur Mutlu |
ISCA | 3 |
| 2011 | Reducing memory interference in multicore systems via application-aware memory channel partitioningabstractMain memory is a major shared resource among cores in a multicore system. If the interference between different applications' memory requests is not controlled effectively, system performance can degrade significantly. Previous work aimed to mitigate the problem of interference between applications by changing the scheduling policy in the memory controller, i.e., by prioritizing memory requests from applications in a way that benefits system performance. Sai Prashanth Muralidhara, Lavanya Subramanian, Onur Mutlu, Mahmut T. Kandemir, Thomas Moscibroda |
MICRO | 2 |