Anup Holey

dblp:117/3141 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 70% GPUs and heterogeneous computing · 23% Processor architecture and microarchitecture · 6%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management › virtual memory
address translation
0.712023
Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding · HPCA 2023
GPUs and heterogeneous computing
multi-GPU computing
0.712023
Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding · HPCA 2023
Memory systems › memory management › virtual memory
page table walk
0.712023
Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding · HPCA 2023
Memory systems › memory management
virtual memory
0.712023
Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding · HPCA 2023
Memory systems
cache management
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache management
cache partitioning
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
Memory systems › cache
shared last-level cache
0.212015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015
GPUs and heterogeneous computing › GPU memory management
unified virtual memory
0.212023
Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding · HPCA 2023
Energy-efficient computing › power management › memory power management
cache energy reduction
0.112015
Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor · ACM Trans. Archit. Code Optim. 2015

Methods — techniques the papers use, named apart from their topics

remote forwarding · 0.7page walk cache · 0.7simulation · 0.2runtime thread-level parallelism measurement · 0.2
YearPublicationVenuePosition
2023 Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote Forwarding
abstract
Multi-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average.
Bingyao Li 0001, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang 0002, Xulong Tang
HPCA3
2015 Performance-Energy Considerations for Shared Cache Management in a Heterogeneous Multicore Processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as graphic processing unit (GPU) cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of concurrent threads supported by the architecture. Under current cache management policies, the CPU applications’ share of the LLC can be significantly reduced in the presence of competing GPU applications. For many CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can tolerate increase in memory access latency when there is sufficient thread-level parallelism (TLP). In addition to the performance challenge, introduction of diverse cores onto the same die changes the energy consumption profile and, in turn, affects the energy efficiency of the processor. In this work, we propose heterogeneous LLC management (HeLM), a novel shared LLC management policy that takes advantage of the GPU’s tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache-sensitive CPU applications. This throttling is achieved by allowing GPU accesses to bypass the LLC when an increase in memory access latency can be tolerated. The latency tolerance of a GPU application is determined by the availability of TLP, which is measured at runtime as the average number of threads that are available for issuing. For a baseline configuration with two CPU cores and four GPU cores, modeled after existing heterogeneous processor designs, HeLM outperforms least recently used (LRU) policy by 10.4%. Additionally, HeLM also outperforms competing policies. Our evaluations show that HeLM is able to sustain performance with varying core mix. In addition to the performance benefit, bypassing also reduces total accesses to the LLC, leading to a reduction in the energy consumption of the LLC module. However, LLC bypassing has the potential to increase off-chip bandwidth utilization and DRAM energy consumption. Our experiments show that HeLM exhibits better energy efficiency by reducing the ED 2 by 18% over LRU while impacting only a 7% increase in off-chip bandwidth utilization.
Anup Holey, Vineeth Mekkat, Pen-Chung Yew, Antonia Zhai
ACM Trans. Archit. Code Optim.1
2014 Lightweight Software Transactions on GPUs
abstract
Graphics Processing Units (GPUs) provide an attractive option for extracting data-level parallelism from diverse applications. However, some applications, although possess abundant data-level parallelism, exhibit irregular memory access patterns to the shared data structures. Porting such applications to GPUs requires synchronization mechanisms such as locks, which significantly increase the programming complexity. Coarse-grained locking, where a single lock controls all the shared resources, although reduces programming efforts, can substantially serialize GPU threads. On the other hand, fine-grained locking, where each data element is protected by an independent lock, although facilitates maximum parallelism, requires significant programming efforts. To overcome these challenges, we propose to support software transactional memory (STM) on GPU that is able to achieve performance comparable to fine-grained locking, while requiring minimal programming efforts. Software-based transactional execution can incur significant runtime overheads due to activities such as detecting conflicts across thousands of GPU threads and managing a consistent memory state. Thus, in this paper we illustrate three lightweight STM designs that are capable of scaling to a large number of GPU threads. In our system, programmers simply mark the critical sections in the applications, and the underlying STM support is able to achieve performance comparable to fine-grained locking.
Anup Holey, Antonia Zhai
ICPP1
2013 Managing shared last-level cache in a heterogeneous multicore processor
abstract
Heterogeneous multicore processors that integrate CPU cores and data-parallel accelerators such as GPU cores onto the same die raise several new issues for sharing various on-chip resources. The shared last-level cache (LLC) is one of the most important shared resources due to its impact on performance. Accesses to the shared LLC in heterogeneous multicore processors can be dominated by the GPU due to the significantly higher number of threads supported. Under current cache management policies, the CPU applications' share of the LLC can be significantly reduced in the presence of competing GPU applications. For cache sensitive CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications can often tolerate increased memory access latency in the presence of LLC misses when there is sufficient thread-level parallelism. In this work, we propose Heterogeneous LLC Management (HeLM), a novel shared LLC management policy that takes advantage of the GPU's tolerance for memory access latency. HeLM is able to throttle GPU LLC accesses and yield LLC space to cache sensitive CPU applications. GPU LLC access throttling is achieved by allowing GPU threads that can tolerate longer memory access latencies to bypass the LLC. The latency tolerance of a GPU application is determined by the availability of thread-level parallelism, which can be measured at runtime as the average number of threads that are available for issuing. Our heterogeneous LLC management scheme outperforms LRU policy by 12.5% and TAP-RRIP by 5.6% for a processor with 4 CPU and 4 GPU cores.
Vineeth Mekkat, Anup Holey, Pen-Chung Yew, Antonia Zhai
PACT2
2013 HAccRG: Hardware-Accelerated Data Race Detection in GPUs
abstract
Modern Graphics Processing Units (GPUs) are capable of supporting thousands of concurrent threads. However, they provide relatively little guarantee with respect to the coherence and consistency of the memory system. Thus, GPUs are prone to multitude of concurrency bugs related to inconsistent memory states. Many such bugs manifest as some form of data races at runtime, and being able to identify these data races can help programmers improve software reliability. Mechanisms that enable efficient and effective data race detection at runtime can form the basis of powerful tools for enhancing GPU software correctness. Most prior works in data race detection for GPU focus on the software-based approaches that incur significant performance overhead. Furthermore, they often focus on the smaller shared memory, while neglecting the larger global memory. We believe that adequate hardware support can enable efficient data race detection in all levels of the memory system for GPUs. In this paper, we propose a hardware-accelerated data race detection mechanism, HAccRG, for efficient data race detection in GPUs. HAccRG provides hardware support for tracking data dependencies across a large number of threads and detects various forms of data races. We incorporate HAccRG on both the shared and global memory spaces in GPU. Our evaluation shows that, with moderate hardware support, HAccRG can detect data races in GPU kernels with a small overhead: 1% for the shared memory and 27% for combined shared and global memory data race detection.
Anup Holey, Vineeth Mekkat, Antonia Zhai
ICPP1
2013 Accelerating Data Race Detection Utilizing On-Chip Data-Parallel Cores
Vineeth Mekkat, Anup Holey, Antonia Zhai
RV2
2012 Energy-efficient non-minimal path on-chip interconnection network for heterogeneous systems
abstract
Network-on-Chips (NoCs) in heterogeneous systems containing both CPU and GPU cores must be designed to satisfy the performance requirements of both latency-sensitive CPU traffic and throughput-intensive GPU traffic. DVFS and adaptive routing can potentially improve NoC energy and performance efficiency. We further notice that GPU traffic can sometimes tolerate a slack defined as the number of cycles a packet can be delayed without causing performance penalty. In this work, we take advantage of the slack in GPU packets to route packets through non-minimal path, so that routers can operate at a lower frequency without suffering performance penalty.
Jieming Yin, Pingqiang Zhou, Anup Holey, Sachin S. Sapatnekar, Antonia Zhai
ISLPED3