EDBT 2026 Demo / reviewers in the wild / expert
John D. Leidel
dblp:129/5531 · also John Leidel
· DBLP profile ↗
12ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-7567-8145ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Application Level Architecture DesignabstractRecent advancements in compiler infrastructures have yielded malleable frameworks to explore data and target optimization using language agnostic approaches. In this talk, we will explore the latest efforts in the compiler and architecture communities to construct frameworks for architecture exploration using traditional language and compiler workflows. John D. Leidel |
CF | 1 |
| 2021 | xBGAS: A Global Address Space Extension on RISC-V for High Performance ComputingabstractThe tremendous expansion of data volume has driven the transition from monolithic architectures towards systems integrated with discrete and distributed subcomponents in modern scalable high performance computing (HPC) systems. As such, multi-layered software infrastructures have become essential to bridge the gap between heterogeneous commodity devices. However, operations across synthesized components with divergent interfaces inevitably lead to redundant software footprints and undesired latency. Therefore, a scalable and unified computing platform, capable of supporting efficient interactions between individual components, is desirable for largescale data-intensive applications. In this work, we introduce the Extended Base Global Address Space, or xBGAS, microarchitecture extension to the RISC-V instruction set architecture (ISA) for scalable high performance computing. The xBGAS extension provides native ISA-level support for direct accesses to remote shared memory by mapping remote data objects into a system's extended address space. We perform both software and hardware evaluations of the xBGAS design. The results show that xBGAS reduces instruction count generated by interprocess communication by 69.26% on average. Overall, xBGAS achieves an average performance gain of 21.96% (up to 37.29%) across the tested workloads. Xi Wang 0009, John D. Leidel, Brody Williams, Alan Ehret, Miguel Mark, Michel A. Kinsy, Yong Chen 0001 |
IPDPS | 2 |
| 2021 | HAM: Hotspot-Aware Manager for Improving Communications With 3D-Stacked MemoryabstractEmerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this article, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51 percent and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81 percent in average (up to 34.28 percent), and power savings of 35.07 percent over a standard 3D-stacked memory. Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001 |
IEEE Trans. Computers | 3 |
| 2020 | StoneCutter: a very high level instruction set design languageabstractAs the density and capability of reconfigurable computing using FPGAs continues to increase and access to large scale ASIC integration continues to increase, research activities associated with high level synthesis flows have expanded at a similar rate. The goal of these research efforts is to reduce the time and effort required to construct and deploy application-specific architectures. However, these synthesis techniques often force users to consider the entire circuit design space in order to develop a successful implementation. This lack of design specificity often results in hardware design implementations that are difficult to program, difficult to reuse in future designs and make sub-optimal use of hardware resources. John D. Leidel, David Donofrio, Frank Conlon |
CF | 1 |
| 2020 | Remote Atomic Extension (RAE) for Scalable High Performance ComputingabstractEmerging data-intensive applications such as graph analytics, machine learning, and data-driven scientific computing are driving the evolution of high-performance computing (HPC) systems from monolithic to scaled-out, heterogeneous, and complex architectures. In these systems, enormous data sets are mapped to discrete nodes to improve the performance of the system by using distributed storage and computing resources. As such, these data distributions induce frequent cross-node data transactions which challenge the performance of large-scale systems. Global atomic operations are one emerging class of the remote data operations that enable lock-free remote shared data operations. However, the cross-node read-modify-write operations consist of multiple distinct data operations and specific atomicity management, which induces a large amount of overhead. As such, these global atomic operations require an efficient communication methodology Existing advanced compo-nents, such as network interface controllers, network fabrics, network-on-chip (NoC) interconnects, are architected together to improve the system performance. However, complex software infrastructures are needed to provide integration between each discrete component. As a result, the redundant software routines across distinct devices induce a large amount of overhead that causes performance degradationIn this paper, we propose a remote atomic extension (RAE) design that provides inherent ISA-level instructions and micro-architecture support for remote atomic operations based on the RISC-V instruction set architecture (ISA). We design a toolchain and evaluate the RAE infrastructure via simulation. Our experiment results show that RAE eliminates 89.71% of the redundant software instructions used for remote atomic accesses and improves the performance by 17.61% on average (up to 23.35%), compared with the OpenSHMEM. Xi Wang 0009, Brody Williams, John D. Leidel, Alan Ehret, Michel A. Kinsy, Yong Chen 0001 |
DAC | 3 |
| 2020 | PAC: Paged Adaptive Coalescer for 3D-Stacked MemoryabstractMany contemporary data-intensive applications exhibit irregular and highly concurrent memory access patterns and thus challenge the performance of conventional memory systems. Driven by an expanding need for high-bandwidth memory featuring low access latency, 3D-stacked memory devices, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), were designed to provide significantly higher throughput as compared to standard JEDEC DDR devices. However, existing memory interfaces and coalescing models, designed for conventional DDR devices, are unable to fully exploit the bandwidth potential inherent in these new 3D-stacked memory devices. In order to remedy this disparity, we introduce in this work a novel paged adaptive coalescer (PAC) infrastructure with a scalable coalescing network for 3D-stacked memory. We present the design and simulated implementation of this approach on RISC-V embedded cores with attached HMC devices. We have carried out extensive evaluations and the results show that the proposed PAC methodology yields an average coalescing efficiency of 56.01%. Further, our evaluation results also show that the PAC reduces bank conflicts and the power consumption by 85.16% and 59.21%, respectively. Overall, PAC achieves an average performance gain of 14.35% (and up to 26.06%) across 14 test suites. These results showcase the potential of the PAC methodology as applied to architecture design for increasingly critical data-intensive algorithms and applications. Xi Wang 0009, John D. Leidel, Brody Williams, Yong Chen 0001 |
HPDC | 2 |
| 2019 | POSTER: Memory Hotspot Optimization for Data-Intensive ApplicationsabstractEmerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, big data science, are data-intensive. The data-intensive workloads usually present irregular memory footprints with limited data locality, and thus incur frequent cache misses and a growing desire for memory bandwidth. Driven by this need, 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) are introduced to yield significantly higher throughput. However, the traditional interfaces and optimization methods for JEDEC DDR devices cannot fully exploit the potential performance of 3D-stacked memory to handle massive irregular memory accesses accompanied with data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices that is capable of optimizing memory access streams via request aggregation, hotspot detection, prefetching, and an associated hotspot-aware page policy. We present the HAM design and simulation implementation on RISC-V embedded cores with attached HMC devices. We have conducted extensive evaluations with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results reveal that HAM reduces redundant memory accesses by 37.51% and achieves a 4.19X enhancement on the prefetch buffer hit rate on average. Overall, HAM exhibits an average of 21.81% performance gain (up to 34.28%) and 35.07% power saving over the standard 3D-stacked memory. Xi Wang 0009, Jie Li 0057, Antonino Tumeo, John D. Leidel, Yong Chen 0001 |
PACT | 4 |
| 2019 | Toward a graph-based dependence analysis framework for high level design verificationabstractRecent efforts to deploy FPGA's and application-specific accelerator devices in scalable data center environments has led to a resurgence in research associated with high level synthesis and design verification. The goal of this research has been to accelerate the initial design, verification and deployment process for abstract accelerator platforms. While the research associated with high level synthesis flows has provided significant gains in design acceleration, research in the verification of these designs has largely been based upon augmenting traditional methodologies. John D. Leidel, Frank Conlon |
CF | 1 |
| 2019 | MAC: Memory Access Coalescer for 3D-Stacked MemoryabstractEmerging data-intensive applications, such as graph analytics and data mining, exhibit irregular memory access patterns. Research has shown that with these memory-bound applications, traditional cache-based processor architectures, which exploit locality and regular patterns to mitigate the memory-wall issue, are inefficient. Meantime, novel 3D-stacked memory devices, such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), promise significant increases in bandwidth that appear extremely appealing for memory-bound applications. However, conventional memory interfaces designed for cache-based architectures and JEDEC DDR devices fit poorly with the 3D-stacked memory, which leads to significant under-utilization of the promised high bandwidth. Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001 |
ICPP | 3 |
| 2018 | Memory Coalescing for Hybrid Memory CubeabstractArguably, many data-intensive applications pose significant challenges to conventional architectures and memory systems, especially when applications exhibit non-contiguous, irregular, and small memory access patterns. The long memory access latency can dramatically slow down the overall performance of applications. The growing desire of high memory bandwidth and low latency access stimulate the advent of novel 3D-staked memory devices such as the Hybrid Memory Cube (HMC), which provides significantly higher bandwidth compared with the conventional JEDEC DDR devices. Even though many existing studies have been devoted to achieving high bandwidth throughput of HMC, the bandwidth potential cannot be fully exploited due to the lack of highly efficient memory coalescing and interfacing methodology for HMC devices. In this research, we introduce a novel memory coalescer methodology that facilitates memory bandwidth efficiency and the overall performance through an efficient and scalable memory request coalescing interface for HMC. We present the design and implementation of this approach on RISC-V embedded cores with attached HMC devices. Our evaluation results show that the new memory coalescer eliminates 47.47% memory accesses to HMC and improves the overall performance by 13.14% on average. Xi Wang 0009, John D. Leidel, Yong Chen 0001 |
ICPP | 2 |
| 2017 | OpenSoC system architect: An open toolkit for building soft-cores on FPGAsabstractGiven the recent difficulty in continuing the classic CMOS manufacturing density and power scaling curves, also known as Moore's Law and Dennard Scaling, respectively, we find that modern complex system architectures are increasingly relying upon accelerators in order to optimize the placement of specific computational workloads. In addition, large-scale computing infrastructures utilized in HPC, data intensive computing, and cloud computing must rely almost exclusively upon commodity device architectures provided by third-party manufacturers. The end result being a final system architecture that lacks specificity for the target software workload. At the same time, there is a trend in the FPGA space of much larger FPGAs with a lot more resources and hardened IP blocks, making this type of architecture design space exploration much easier. The OpenSoC System Architect infrastructure combines several open source design tools and methodologies into a central infrastructure for designing, developing, and verifying the necessary hardware and software modules required to implement application-specific processors for use in FPGAs. The end result is an infrastructure that permits rapid development and deployment of application-specific accelerators and softcores, including a fully functional software development tool chain. Farzad Fatollahi-Fard, David Donofrio, John Shalf, John D. Leidel, Xi Wang 0009, Yong Chen 0001 |
FPL | 4 |
| 2017 | HMC-Sim-2.0: A co-design infrastructure for exploring custom memory cube operations
John D. Leidel, Yong Chen 0001 |
Parallel Comput. | 1 |