VLDB 2026 Research / reviewers in the wild / expert
Michael LeBeane
dblp:156/2976
· DBLP profile ↗
12ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 35% GPUs and heterogeneous computing · 34% Parallel and multicore computing · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 18 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory management › virtual memory
address translation |
0.8 | 2 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 Neighborhood-Aware Address Translation for Irregular GPU Applications · MICRO 2018 |
GPUs and heterogeneous computing
GPU memory management |
0.8 | 2 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 Neighborhood-Aware Address Translation for Irregular GPU Applications · MICRO 2018 |
Memory systems › memory management
virtual memory |
0.8 | 2 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 Neighborhood-Aware Address Translation for Irregular GPU Applications · MICRO 2018 |
GPUs and heterogeneous computing › GPU communication
GPU networking |
0.7 | 2 | 2020 | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs · PPoPP 2020 GPU triggered networking for intra-kernel communications · SC 2017 |
GPUs and heterogeneous computing › GPU memory management
GPU address translation |
0.5 | 1 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 |
Memory systems › memory management › virtual memory
TLB reach |
0.5 | 1 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 |
Distributed systems › distributed communication
remote memory access |
0.4 | 1 | 2020 | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs · PPoPP 2020 |
Memory systems › virtual memory management
address translation overhead |
0.3 | 1 | 2018 | Neighborhood-Aware Address Translation for Irregular GPU Applications · MICRO 2018 |
Parallel and multicore computing › parallel programming runtimes
active messages |
0.2 | 1 | 2016 | Extended task queuing: active messages for heterogeneous systems · SC 2016 |
Interconnection networks and networks-on-chip
remote direct memory access |
0.2 | 1 | 2016 | Extended task queuing: active messages for heterogeneous systems · SC 2016 |
Parallel and multicore computing
data distribution |
0.2 | 1 | 2015 | Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015 |
Parallel and multicore computing
graph processing |
0.2 | 1 | 2015 | Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015 |
High-performance computing › cluster computing
heterogeneous clusters |
0.2 | 1 | 2015 | Data partitioning strategies for graph workloads on heterogeneous clusters · SC 2015 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.1 | 1 | 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources · MICRO 2021 |
Interconnection networks and networks-on-chip
network interface |
0.1 | 1 | 2020 | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUs · PPoPP 2020 |
Memory systems › memory management › virtual memory
page table |
0.1 | 1 | 2018 | Neighborhood-Aware Address Translation for Irregular GPU Applications · MICRO 2018 |
Parallel and multicore computing › parallel computing › parallel communication
inter-node communication |
0.1 | 1 | 2017 | GPU triggered networking for intra-kernel communications · SC 2017 |
High-performance computing › collective communication
MPI collective communication |
0.1 | 1 | 2016 | Extended task queuing: active messages for heterogeneous systems · SC 2016 |
Methods — techniques the papers use, named apart from their topics
microarchitecture simulation · 0.7instruction-level analysis · 0.7victim cache · 0.5on-chip resource reuse · 0.5neighborhood-aware coalescing · 0.3cache line reuse · 0.3GPU-triggered networking · 0.3task queuing · 0.2active messages · 0.2RDMA · 0.2heterogeneity-aware data ingress · 0.2graph cutting · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip ResourcesabstractMany GPU applications issue irregular memory accesses to a very large memory footprint. We confirm observations from prior work that these irregular access patterns are severely bottlenecked by insufficient Translation Lookaside Buffer (TLB) reach, resulting in expensive page table walks. In this work, we investigate mechanisms to improve TLB reach without increasing the page size or the size of the TLB itself. Our work is based around the observation that a GPU’s instruction cache (I-cache) and Local Data Share (LDS) scratchpad memory are under-utilized in many applications, including those that suffer from poor TLB reach. We leverage this to opportunistically utilize idle capacity and port bandwidth from the GPU’s I-cache and LDS structures for address translations. We explore various potential architectural designs for each structure to optimize performance and minimize complexity. Both structures are organized as a victim cache between the L1 and L2 TLBs to boost translation reach. We find that our designs can increase performance on average by 30.1% without impacting the performance of applications that do not require additional reach. Jagadish Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. Loh |
MICRO | 2 |
| 2020 | <u>G</u>PU <u>i</u>nitiated <u>O</u>penSHMEM: correct and efficient intra-kernel networking for dGPUsabstractCurrent state-of-the-art in GPU networking utilizes a host-centric, kernel-boundary communication model that reduces performance and increases code complexity. To address these concerns, recent works have explored performing network operations from within a GPU kernel itself. However, these approaches typically involve the CPU in the critical path, which leads to high latency and inefficient utilization of network and/or GPU resources. Khaled Hamidouche, Michael LeBeane |
PPoPP | 2 |
| 2018 | ComP-net: command processor networking for efficient intra-kernel communications on GPUsabstractCurrent state-of-the-art in GPU networking advocates a host-centric model that reduces performance and increases code complexity. Recently, researchers have explored several techniques for networking within a GPU kernel itself. These approaches, however, suffer from high latency, waste energy on the host, and are not scalable with larger/more GPUs on a node. In this work, we introduce Command Processor Networking (ComP-Net), which leverages the availability of scalar cores integrated on the GPU itself to provide high-performance intra-kernel networking. ComP-Net enables efficient synchronization between the Command Processors and Compute Units on the GPU through a line locking scheme implemented in the GPU's shared last-level cache. We illustrate that ComP-Net can improve application performance by up to 20% and provide up to 50% reduction in energy consumption vs. competing networking techniques across a Jacobi stencil, allreduce collective, and machine learning applications. Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John |
PACT | 1 |
| 2018 | Lost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediate Language LevelabstractModern GPU frameworks use a two-phase compilation approach. Kernels written in a high-level language are initially compiled to an implementation agnostic intermediate language (IL), then finalized to the machine ISA only when the target GPU hardware is known. Most GPU microarchitecture simulators available to academics execute IL instructions because there is substantially less functional state associated with the instructions, and in some situations, the machine ISA's intellectual property may not be publicly disclosed. In this paper, we demonstrate the pitfalls of evaluating GPUs using this higher-level abstraction, and make the case that several important microarchitecture interactions are only visible when executing lower-level instructions. Our analysis shows that given identical application source code and GPU microarchitecture models, execution behavior will differ significantly depending on the instruction set abstraction. For example, our analysis shows the dynamic instruction count of the machine ISA is nearly 2× that of the IL on average, but contention for vector registers is reduced by 3× due to the optimized resource utilization. In addition, our analysis highlights the deficiencies of using IL to model instruction fetching, control divergence, and value similarity. Finally, we show that simulating IL instructions adds 33% error as compared to the machine ISA when comparing absolute runtimes to real hardware. Anthony Gutierrez, Bradford M. Beckmann, Alexandru Dutu, Joseph Gross, Michael LeBeane, John Kalamatianos, Onur Kayiran, Matthew Poremba, Brandon Potter, Sooraj Puthoor, Matthew D. Sinclair, Mark Wyse, Jieming Yin, Xianwei Zhang 0001, Akshay Jain 0005, Timothy G. Rogers |
HPCA | 5 |
| 2018 | Neighborhood-Aware Address Translation for Irregular GPU ApplicationsabstractRecent studies on commercial hardware demonstrated that irregular GPU workloads could bottleneck on virtual-to-physical address translations. GPU's single-instruction multiple-thread (SIMT) execution can generate many concurrent memory accesses, all of which require address translation before accesses can complete. Unfortunately, many of these address translation requests often miss in the TLB, generating many concurrent page table walks. In this work, we investigate how to reduce address translation overheads for such applications. We observe that many of these concurrent page walk requests, while irregular from the perspective of a single GPU wavefront, still fall on neighboring virtual page addresses. The address mappings for these neighboring pages are typically stored in the same 64-byte cache line. Since cache lines are the smallest granularity of memory access, the page table walker implicitly reads address mappings (i.e., page table entries or PTEs) of many neighboring pages during the page walk of a single virtual address (VA). However, in the conventional hardware, mappings not associated with the original request are simply discarded. In this work, we propose mechanisms to coalesce the address translation needs of all pending page table walks in the same neighborhood that happen to have their address mappings fall on the same cache line. This is almost free; the page table walker (PTW) already reads a full cache line containing address mappings of all pages in the same neighborhood. We find this simple scheme can reduce the number of accesses to the inmemory page table by 37% on average. This speeds up a set of GPU workloads by an average of 1.7×. Seunghee Shin, Michael LeBeane, Yan Solihin, Arkaprava Basu |
MICRO | 2 |
| 2017 | GPU triggered networking for intra-kernel communicationsabstractGPUs are widespread across clusters of compute nodes due to their attractive performance for data parallel codes. However, communicating between GPUs across the cluster is cumbersome when compared to CPU networking implementations. A number of recent works have enabled GPUs to more naturally access the network, but suffer from performance problems, require hidden CPU helper threads, or restrict communications to kernel boundaries. Michael LeBeane, Khaled Hamidouche, Brad Benton, Maurício Breternitz, Steven K. Reinhardt, Lizy Kurian John |
SC | 1 |
| 2016 | Proxy-Guided Load Balancing of Graph Processing Workloads on Heterogeneous ClustersabstractBig data decision-making techniques take advantage of large-scale data to extract important insights from them. One of the most important classes of such techniques falls in the domain of graph applications, where data segments and their inherent relationships are represented as vertices and edges. Efficiently processing large-scale graphs involves many subtle tradeoffs and is still regarded as an open-ended problem. Furthermore, as modern data centers move towards increased heterogeneity, the traditional assumption of homogeneous environments in current graph processing frameworks is no longer valid. Prior work estimates the graph processing power of heterogeneous machines by simply reading hardware configurations, which leads to suboptimal load balancing. In this paper, we propose a profiling methodology leveraging synthetic graphs for capturing a node's computational capability and guiding graph partitioning in heterogeneous environments with minimal overheads. We show that by sampling the execution of applications on synthetic graphs following a power-law distribution, the computing capabilities of heterogeneous clusters can be captured accurately (<;10% error). Our proxy-guided graph processing system results in a maximum speedup of 1.84x and 1.45x over a default system and prior work, respectively. On average, it achieves 17.9% performance improvement and 14.6% energy reduction as compared to prior heterogeneity-aware work. Shuang Song 0007, Xinnian Zheng, Michael LeBeane, Jeeho Ryoo, Reena Panda, Andreas Gerstlauer, Lizy Kurian John |
ICPP | 4 |
| 2016 | Extended task queuing: active messages for heterogeneous systemsabstractAccelerators have emerged as an important component of modern cloud, datacenter, and HPC computing environments. However, launching tasks on remote accelerators across a network remains unwieldy, forcing programmers to send data in large chunks to amortize the transfer and launch overhead. By combining advances in intra-node accelerator unification with one-sided Remote Direct Memory Access (RDMA) communication primitives, it is possible to efficiently implement lightweight tasking across distributed-memory systems. This paper introduces Extended Task Queuing (XTQ), an RDMA-based active messaging mechanism for accelerators in distributed-memory systems. XTQ's direct NIC-to-accelerator communication decreases inter-node GPU task launch latency by 10-15% for small-to-medium sized messages and ameliorates CPU message servicing overheads. These benefits are shown in the context of MPI accumulate, reduce, and allreduce operations with up to 64 nodes. Finally, we illustrate how XTQ can improve the performance of popular deep learning workloads implemented in the Computational Network Toolkit (CNTK). Michael LeBeane, Brandon Potter, Abhisek Pan, Alexandru Dutu, Vinay Agarwala, Wonchan Lee, Deepak Majeti, Bibek Ghimire, Eric Van Tassell, Samuel Wasmundt, Brad Benton, Maurício Breternitz, Michael L. Chu, Mithuna Thottethodi, Lizy Kurian John, Steven K. Reinhardt |
SC | 1 |
| 2015 | GPGPU Benchmark Suites: How Well Do They Sample the Performance Spectrum?abstractRecently, GPGPUs have positioned themselves in the mainstream processor arena with their potential to perform a massive number of jobs in parallel. At the same time, many GPGPU benchmark suites have been proposed to evaluate the performance of GPGPUs. Both academia and industry have been introducing new sets of benchmarks each year while some already published benchmarks have been updated periodically. However, some benchmark suites contain benchmarks that are duplicates of each other or use the same underlying algorithm. This results in an excess of workloads in the same performance spectrum. In this paper, we provide a methodology to obtain a set of new GPGPU benchmarks that are located in the unexplored region of the performance spectrum. Our proposal uses statistical methods to understand the performance spectrum coverage and uniqueness of existing benchmark suites. Later we show techniques to identify areas that are not explored by existing benchmarks by visually showing the performance spectrum coverage. Finding unique key metrics for future benchmarks to broaden its performance spectrum coverage is also explored using hierarchical clustering and ranking by Hotel ling's T2 method. Finally, key metrics are categorized into GPGPU performance related components to show how future benchmarks can stress each of the categorized metrics to distinguish themselves in the performance spectrum. Our methodology can serve as a performance spectrum oriented guidebook for designing future GPGPU benchmarks. Jeeho Ryoo, Saddam Quirem, Michael LeBeane, Reena Panda, Shuang Song 0007, Lizy Kurian John |
ICPP | 3 |
| 2015 | Watt Watcher: Fine-Grained Power Estimation for Emerging WorkloadsabstractExtensive research has focused on estimating power to guide advances in power management schemes, thermal hot spots, and voltage noise. However, simulated power models are slow and struggle with deep software stacks, while direct measurements are typically coarse-grained. This paper introduces Watt Watcher, a multicore power measurement framework that offers fine-grained functional unit breakdowns. Watt Watcher operates by passing event counts and a hardware descriptor file into configurable back-end power models based on McPAT. Researchers and vendors can add other processors to our tool by mapping to the Watt Watcher interface. We show that Watt Watcher, when calibrated, has a MAPE (mean absolute percentage error) of 2.67% aggregated over all benchmarks when compared to measured power consumption on SPEC CPU 2006 and multithreaded PARSEC benchmarks across three different machines of various form factors and manufacturing processes. We present two use cases showing how Watt Watcher can derive insights that are difficult to obtain through other measurement infrastructures. Additionally, we illustrate how Watt Watcher can be used to provide insights into challenging big data and cloud workloads on a server CPU. Through the use of Watt Watcher, it is possible to obtain a detailed power breakdown on real hardware without vendor proprietary models or hardware instrumentation. Michael LeBeane, Jeeho Ryoo, Reena Panda, Lizy Kurian John |
SBAC-PAD | 1 |
| 2015 | Performance Characterization of Modern Databases on Out-of-Order CPUsabstractBig data revolution has created an unprecedented demand for intelligent data management solutions on a large scale. While data management has traditionally been used as a synonym for relational data processing, in recent years a new group popularly known as NoSQL databases have emerged as a competitive alternative. There is a pressing need to gain greater understanding of the characteristics of modern databases to architect targeted computers. In this paper, we investigate four popular NoSQL/SQL-style databases and evaluate their hardware performance on modern computer systems. Based on data collected from real hardware, we evaluate how efficiently modern databases utilize the underlying systems and make several recommendations to improve their performance efficiency. We observe that performance of modern databases is severely limited by poor cache/memory performance. Nonetheless, we demonstrate that dynamic execution techniques are still effective in hiding a significant fraction of the stalls, thereby improving performance. We further show that NoSQL databases suffer from greater performance inefficiencies than their SQL counterparts. SQL databases outperform NoSQL databases for most operations and are beaten by NoSQL databases only in a few cases. NoSQL databases provide a promising competitive alternative to SQL-style databases, however, they are yet to be optimized to fully reach the performance of contemporary SQL systems. We also show that significant diversity exists among different database implementations and big-data benchmark designers can leverage our analysis to incorporate representative workloads to encapsulate the full spectrum of data-serving applications. In this paper, we also compare data-serving applications with other popular benchmarks such as SPEC CPU2006 and SPECjbb2005. Reena Panda, Christopher Erb, Michael LeBeane, Jeeho Ryoo, Lizy Kurian John |
SBAC-PAD | 3 |
| 2015 | Data partitioning strategies for graph workloads on heterogeneous clustersabstractLarge scale graph analytics are an important class of problem in the modern data center. However, while data centers are trending towards a large number of heterogeneous processing nodes, graph analytics frameworks still operate under the assumption of uniform compute resources. In this paper, we develop heterogeneity-aware data ingress strategies for graph analytics workloads using the popular PowerGraph framework. We illustrate how simple estimates of relative node computational throughput can guide heterogeneity-aware data partitioning algorithms to provide balanced graph cutting decisions. Our work enhances five online data ingress strategies from a variety of sources to optimize application execution for throughput differences in heterogeneous data centers. The proposed partitioning algorithms improve the runtime of several popular machine learning and data mining applications by as much as a 65% and on average by 32% as compared to the default, balanced partitioning approaches. Michael LeBeane, Shuang Song 0007, Reena Panda, Jeeho Ryoo, Lizy Kurian John |
SC | 1 |