Vineeth Mekkat

dblp:115/5253 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 70% Processor architecture and microarchitecture · 23% Energy-efficient computing · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache management
cache partitioning
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache
shared last-level cache
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › power management › memory power management
cache energy reduction
0.112015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015

Methods — techniques the papers use, named apart from their topics

simulation · 0.2runtime thread-level parallelism measurement · 0.2
YearPublicationVenuePosition
2015 Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as graphic processing unit (GPU) cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of concurrent threads supported by the architecture. Under current cache management policies, the CPU applications’ share of the LLC can be significantly reduced in the presence of competing GPU applications. For many CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can tolerate increase in memory access latency when there is sufficient thread-level parallelism (TLP). In addition to the performance challenge, introduction of diverse cores onto the same die changes the energy consumption profile and, in turn, affects the energy efficiency of the processor. In this work, we propose heterogeneous LLC management (HeLM), a novel shared LLC management policy that takes advantage of the GPU’s tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache-sensitive CPU applications. This throttling is achieved by allowing GPU accesses to bypass the LLC when an increase in memory access latency can be tolerated. The latency tolerance of a GPU application is determined by the availability of TLP, which is measured at runtime as the average number of threads that are available for issuing. For a baseline configuration with two CPU cores and four GPU cores, modeled after existing heterogeneous processor designs, HeLM outperforms least recently used (LRU) policy by 10.4%. Additionally, HeLM also outperforms competing policies. Our evaluations show that HeLM is able to sustain performance with varying core mix. In addition to the performance benefit, bypassing also reduces total accesses to the LLC, leading to a reduction in the energy consumption of the LLC module. However, LLC bypassing has the potential to increase off-chip bandwidth utilization and DRAM energy consumption. Our experiments show that HeLM exhibits better energy efficiency by reducing the ED 2 by 18% over LRU while impacting only a 7% increase in off-chip bandwidth utilization.
Anup Holey, Vineeth Mekkat, Pen-Chung Yew, Antonia Zhai
ACM Trans. Archit. Code Optim.2
2013 Managing shared last-level cache in a heterogeneous multicore processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as GPU cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of threads supported. Under current cache management policies, the CPU applications' share of the LLC can be significantly reduced in the presence of competing GPU applications. For cache sensitive CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can often tolerate increased memory access latency in the presence of LLC misses when there is sufficient thread-level parallelism. In this work, we propose Heterogeneous LLC Management (HeLM), a novel shared LLC management policy that takes advantage of the GPU's tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache sensitive CPU applications. GPU LLC access throttling is achieved by allowing GPU threads that can tolerate longer memory access latencies to bypass the LLC. The latency tolerance of a GPU application is determined by the availability of thread-level parallelism, which can be measured at runtime as the average number of threads that are available for issuing. Our heterogeneous LLC management scheme outperforms LRU policy by 12.5% and TAP-RRIP by 5.6% for a processor with 4 CPU and 4 GPU cores.
Vineeth Mekkat, Anup Holey, Pen-Chung Yew, Antonia Zhai
PACT1
2013 HAccRG: Hardware-Accelerated Data Race Detection in GPUs
abstract
Modern Graphics Processing Units (GPUs) are capable of supporting thousands of concurrent threads. However, they provide relatively little guarantee with respect to the coherence and consistency of the memory system. Thus, GPUs are prone to multitude of concurrency bugs related to inconsistent memory states. Many such bugs manifest as some form of data races at runtime, and being able to identify these data races can help programmers improve software reliability. Mechanisms that enable efficient and effective data race detection at runtime can form the basis of powerful tools for enhancing GPU software correctness. Most prior works in data race detection for GPU focus on the software-based approaches that incur significant performance overhead. Furthermore, they often focus on the smaller shared memory, while neglecting the larger global memory. We believe that adequate hardware support can enable efficient data race detection in all levels of the memory system for GPUs. In this paper, we propose a hardware-accelerated data race detection mechanism, HAccRG, for efficient data race detection in GPUs. HAccRG provides hardware support for tracking data dependencies across a large number of threads and detects various forms of data races. We incorporate HAccRG on both the shared and global memory spaces in GPU. Our evaluation shows that, with moderate hardware support, HAccRG can detect data races in GPU kernels with a small overhead: 1% for the shared memory and 27% for combined shared and global memory data race detection.
Anup Holey, Vineeth Mekkat, Antonia Zhai
ICPP2
2013 Accelerating Data Race Detection Utilizing On-Chip Data-Parallel Cores
Vineeth Mekkat, Anup Holey, Antonia Zhai
RV1