Hongwen Dai

dblp:147/4962 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 52% Memory systems · 35% Performance modeling and evaluation · 9%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU scheduling
concurrent kernel execution
0.722019
Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution · ACM Trans. Archit. Code Optim. 2019
Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls · HPCA 2018
Memory systems › memory bandwidth management
bandwidth partitioning
0.412019
Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution · ACM Trans. Archit. Code Optim. 2019
GPUs and heterogeneous computing
GPU resource management
0.412019
Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution · ACM Trans. Archit. Code Optim. 2019
Memory systems
memory stall
0.312018
Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls · HPCA 2018
Memory systems › cache management › cache insertion policy
cache bypassing
0.212016
A model-driven approach to warp/thread-block level GPU cache bypassing · DAC 2016
Performance modeling and evaluation › cache performance modeling
cache contention modeling
0.212016
A model-driven approach to warp/thread-block level GPU cache bypassing · DAC 2016
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.212016
A model-driven approach to warp/thread-block level GPU cache bypassing · DAC 2016

Methods — techniques the papers use, named apart from their topics

dominant resource fairness · 0.4CTA combination · 0.4memory access balancing · 0.3inflight memory instruction limiting · 0.3cache partitioning · 0.3performance modeling · 0.2
YearPublicationVenuePosition
2019 Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution
abstract
Contemporary GPUs support multiple kernels to run concurrently on the same streaming multiprocessors (SMs). Recent studies have demonstrated that such concurrent kernel execution (CKE) improves both resource utilization and computational throughput. Most of the prior works focus on partitioning the GPU resources at the cooperative thread array (CTA) level or the warp scheduler level to improve CKE. However, significant performance slowdown and unfairness are observed when latency-sensitive kernels co-run with bandwidth-intensive ones. The reason is that bandwidth over-subscription from bandwidth-intensive kernels leads to much aggravated memory access latency, which is highly detrimental to latency-sensitive kernels. Even among bandwidth-intensive kernels, more intensive kernels may unfairly consume much higher bandwidth than less-intensive ones. In this article, we first make a case that such problems cannot be sufficiently solved by managing CTA combinations alone and reveal the fundamental reasons. Then, we propose a coordinated approach for CTA combination and bandwidth partitioning. Our approach dynamically detects co-running kernels as latency sensitive or bandwidth intensive. As both the DRAM bandwidth and L2-to-L1 Network-on-Chip (NoC) bandwidth can be the critical resource, our approach partitions both bandwidth resources coordinately along with selecting proper CTA combinations. The key objective is to allocate more CTA resources for latency-sensitive kernels and more NoC/DRAM bandwidth resources to NoC-/DRAM-intensive kernels. We achieve it using a variation of dominant resource fairness (DRF). Compared with two state-of-the-art CKE optimization schemes, SMK [52] and WS [55], our approach improves the average harmonic speedup by 78% and 39%, respectively. Even compared to the best possible CTA combinations, which are obtained from an exhaustive search among all possible CTA combinations, our approach improves the harmonic speedup by up to 51% and 11% on average.
Hongwen Dai, Mike Mantor, Huiyang Zhou
ACM Trans. Archit. Code Optim.2
2018 Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls
abstract
Following the advances in technology scaling, graphics processing units (GPUs) incorporate an increasing amount of computing resources and it becomes difficult for a single GPU kernel to fully utilize the vast GPU resources. One solution to improve resource utilization is concurrent kernel execution (CKE). Early CKE mainly targets the leftover resources. However, it fails to optimize the resource utilization and does not provide fairness among concurrent kernels. Spatial multitasking assigns a subset of streaming multiprocessors (SMs) to each kernel. Although achieving better fairness, the resource underutilization within an SM is not addressed. Thus, intra-SM sharing has been proposed to issue thread blocks from different kernels to each SM. However, as shown in this study, the overall performance may be undermined in the intra-SM sharing schemes due to the severe interference among kernels. Specifically, as concurrent kernels share the memory subsystem, one kernel, even as computing-intensive, may starve from not being able to issue memory instructions in time. Besides, severe L1 D-cache thrashing and memory pipeline stalls caused by one kernel, especially a memory-intensive one, will impact other kernels, further hurting the overall performance. In this study, we investigate various approaches to overcome the aforementioned problems exposed in intra-SM sharing. We first highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose two approaches to reduce memory pipeline stalls. The first is to balance memory accesses of concurrent kernels. The second is to limit the number of inflight memory instructions issued from individual kernels. Our evaluation shows that the proposed schemes significantly improve the weighted speedup of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK, by 24.6% and 27.2% on average, respectively, with lightweight hardware overhead.
Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou
HPCA1
2017 POSTER: Accelerate GPU Concurrent Kernel Execution by Mitigating Memory Pipeline Stalls
abstract
In this study, we demonstrate that the performance may be undermined in the state-of-the-art intra-SM sharing schemes for concurrent kernel execution (CKE) on GPUs, due to the interference among concurrent kernels. We highlight that cache partitioning techniques proposed for CPUs are not effective for GPUs. Then we propose to balance memory accesses and limit the number of inflight memory instructions issued from concurrent kernels to reduce memory pipeline stalls. Our proposed schemes significantly improve the performance of two state-of-the-art intra-SM sharing schemes, Warped-Slicer and SMK.
Hongwen Dai, Chao Li 0004, Chen Zhao 0009, Fei Wang 0008, Nanning Zheng 0001, Huiyang Zhou
PACT1
2016 A model-driven approach to warp/thread-block level GPU cache bypassing
abstract
The high amount of memory requests from massive threads may easily cause cache contention and cache-miss-related resource congestion on GPUs. This paper proposes a simple yet effective performance model to estimate the impact of cache contention and resource congestion as a function of the number of warps/thread blocks (TBs) to bypass the cache. Then we design a hardware-based dynamic warp/thread-block level GPU cache bypassing scheme, which achieves 1.68x speedup on average on a set of memory-intensive benchmarks over the baseline. Compared to prior works, our scheme achieves 21.6% performance improvement over SWL-best [29] and 11.9% over CBWT-best [4] on average.
Hongwen Dai, Chao Li 0004, Huiyang Zhou, Saurabh Gupta 0002, Christos Kartsaklis, Mike Mantor
DAC1
2015 Locality-Driven Dynamic GPU Cache Bypassing
abstract
This paper presents novel cache optimizations for massively parallel, throughput-oriented architectures like GPUs. L1 data caches (L1 D-caches) are critical resources for providing high-bandwidth and low-latency data accesses. However, the high number of simultaneous requests from single-instruction multiple-thread (SIMT) cores makes the limited capacity of L1 D-caches a performance and energy bottleneck, especially for memory-intensive applications. We observe that the memory access streams to L1 D-caches for many applications contain a significant amount of requests with low reuse, which greatly reduce the cache efficacy. Existing GPU cache management schemes are either based on conditional/reactive solutions or hit-rate based designs specifically developed for CPU last level caches, which can limit overall performance.
Chao Li 0004, Shuaiwen Song, Hongwen Dai, Albert Sidelnik, Siva Kumar Sastry Hari, Huiyang Zhou
ICS3
2015 Analyzing graphics processor unit (GPU) instruction set architectures
abstract
Because of their high throughput and power efficiency, massively parallel architectures like graphics processing units (GPUs) become a popular platform for generous purpose computing. However, there are few studies and analyses on GPU instruction set architectures (ISAs) although it is wellknown that the ISA is a fundamental design issue of all modern processors including GPUs.
Kothiya Mayank, Hongwen Dai, Jizeng Wei, Huiyang Zhou
ISPASS2
2014 Understanding the tradeoffs between software-managed vs. hardware-managed caches in GPUs
abstract
On-chip caches are commonly used in computer systems to hide long off-chip memory access latencies. To manage on-chip caches, either software-managed or hardware-managed schemes can be employed. State-of-art accelerators, such as the NVIDIA Fermi or Kepler GPUs and Intel's forthcoming MIC “Knights Landing” (KNL), support both software-managed caches, aka. shared memory (GPUs) or near memory (KNL), and hardware-managed L1 data caches (D-caches). Furthermore, shared memory and the L1 D-cache on a GPU utilize the same physical storage and their capacity can be configured at runtime (same for KNL). In this paper, we present an in-depth study to reveal interesting and sometimes unexpected tradeoffs between shared memory and the hardware-managed L1 D- caches in GPU architecture. In our study, the kernels utilizing the L1 D-caches are generated from those leveraging shared memory to ensure that the same optimizations such as tiling are applied equally in both versions. Our detailed analyses reveal that rather than cache hit rates, the following tradeoffs often have more profound performance impacts. On one hand, the kernels utilizing the L1 caches may support higher degrees of thread-level parallelism, offer more opportunities for data to be allocated in registers, and sometimes result in lower dynamic instruction counts. On the other hand, the applications utilizing shared memory enable more coalesced accesses and tend to achieve higher degrees of memory-level parallelism. Overall, our results show that most benchmarks perform significantly better with shared memory than the L1 D-caches due to the high impact of memory-level parallelism and memory coalescing.
Chao Li 0004, Yi Yang 0018, Hongwen Dai, Shengen Yan, Frank Mueller 0001, Huiyang Zhou
ISPASS3