EDBT 2026 Demo / reviewers in the wild / expert
Seyong Lee
dblp:43/1488
· DBLP profile ↗
39ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0001-8872-4932ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 11 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accl++ : A high-productivity programming language for performance and code portability on heterogeneous systemsabstractThis work describes the Accl++ programming language for heterogeneous computing. Accl++ is embedded in the C++ language and implemented as a C++ library. The language allows for describing device code and execution libraries and includes primitives for runtime compilation (RTC), thereby enabling code portability across disparate devices. Here, we demonstrate Accl++’s capability by coding different benchmarks from different domains. This work also describes an analysis of the overheads introduced by the Accl++ RTC support as well as the performance of the generated code for the Accl++ kernels. Accl++ improves heterogeneous code portability without incurring high levels of overhead. In two different heterogeneous systems—one composed of two 32-core AMD EPYC 7513 CPUs and two NVIDIA A100 GPUs and the other composed of two 12-core AMD EPYC 7272 CPUs and two AMD MI100 GPUs—Accl++ enables the execution of the same binary application with observed RTC overheads in the range of 3%–10% of the total kernel execution time, resulting in performance levels similar to those of native CUDA/HIP/OpenCL code. Marc González 0001, Pedro Valero-Lara, Mohammad Alaul Haque Monil, Seyong Lee, Beau Johnston, Aaron R. Young, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter |
Future Gener. Comput. Syst. | 4 |
| 2024 | CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous SystemsabstractPerformance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems. Norihisa Fujita, Beau Johnston, Narasinga Rao Miniskar, Ryohei Kobayashi 0001, Mohammad Alaul Haque Monil, Keita Teranishi, Seyong Lee, Jeffrey S. Vetter, Taisuke Boku |
e-Science | 7 |
| 2024 | sKokkos: Enabling Kokkos with Transparent Device Selection on Heterogeneous Systems using OpenACCabstractThis paper presents a new feature to enable Kokkos with transparent device selection. For application developers, it is not easy to identify which device is the most appropriate to use in a heterogeneous system, since this depends on the characteristics of both the application and the hardware. In Kokkos, a backend is associated with one specific programming model/hardware. Programmers decide which backend to use at compilation time. This new feature implemented on the OpenACC backend eliminates the burden of deciding which device to use, providing a highly productive programming solution for Kokkos applications. This work includes implementation details and a performance study conducted with a set of mini-benchmarks (i.e., AXPY and dot product), kernels (Lattice-Bolzmann method), and two mini-apps (LULESH and miniFE) on two heterogeneous systems with different hardware capabilities. This new Kokkos feature provides high accelerations of up to 35 × thanks to automatic and transparent device selection. Pedro Valero-Lara, Seyong Lee, Joel E. Denny, Keita Teranishi, Jeffrey S. Vetter, Marc González 0001 |
HPC Asia | 2 |
| 2024 | IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous ComputingabstractFrom edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead. Seyong Lee, Beau Johnston, Jeffrey S. Vetter |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell TrajectoryabstractThe ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport. Daniel F. Puleri, Sayan Roychowdhury, Peter Balogh, John Gounley, Erik W. Draeger, Jeff Ames, Adebayo Adebiyi, Simbarashe Chidyagwai, Benjamín Hernández, Seyong Lee, Shirley V. Moore, Jeffrey S. Vetter, Amanda Randles |
CLUSTER | 10 |
| 2021 | Optimization with the OpenACC-to-FPGA framework on the Arria 10 and Stratix 10 FPGAs
Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
Parallel Comput. | 2 |
| 2021 | Analysis of GPU Data Access Patterns on Complex Geometries for the D3Q19 Lattice Boltzmann AlgorithmabstractGPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When implemented with GPUs on complex domains, typically, geometric data is accessed indirectly and lattice data is accessed lexicographically. Although there are a variety of other options, no study has examined the relative efficacy between them. Here, we examine a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct is often better suited than the more common indirect addressing, providing increased computational speed and reducing memory consumption. For the layout, we find that the Collected Structure of Arrays (CSoA) and bundling layouts outperform the common Structure of Array layout; on V100 and P100 devices, CSoA consistently outperforms bundling, however the relationship is more complicated on K40 devices. When compared to state-of-the-art practices, our recommendations lead to speedups of 10-40 percent and reduce memory consumption up to 17 percent. Using performance modeling and computational experimentation, we determine the mechanisms behind the accelerations. We demonstrate that our results hold across multiple GPUs on two leadership class systems, and present the first near-optimal strong results for LBM with arterial geometries run on GPUs. Gregory Herschlag, Seyong Lee, Jeffrey S. Vetter, Amanda Randles |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-OffsabstractIntegrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed. Mohammad Alaul Haque Monil, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
PACT | 3 |
| 2020 | Cash: A Single-Source Hardware-Software Codesign Framework for Rapid PrototypingabstractWith Moore's Law coming to an end, hardware specialization and systems on chips are providing new opportunities for continuing performance scaling while reducing the energy cost of computation. However, the current hardware design methodologies require significant engineering efforts and domain expertise, making the design process unscalable. More importantly, hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and design efficiency. In this work, we introduce Cash, a single-source hardware-software co-design framework for rapid SoC prototyping and accelerators research. Cash leverages the unique efficiency and generative attributes of Modern C++ to provide a unified development environment, aiming at closing the architecture research methodology gap. The Cash framework introduces new co-design programming abstractions that enable seamless integration with existing software from architecture research simulators to high-level synthesis. Blaise-Pascal Tine, Fares Elsabbagh, Seyong Lee, Jeffrey S. Vetter, Hyesoon Kim |
FPGA | 3 |
| 2020 | Productive Hardware Designs using Hybrid HLS-RTL DevelopmentabstractCurrent High-Level Synthesis frameworks provide a productive hardware development methodology where hardware accelerators are generated directly from high-level languages like C/C++ or OpenCL, allowing software developers to quickly accelerate their applications. However, the hardware generated by these frameworks is sub-optimal compared to often hand-optimized RTL modules. A hybrid development approach would leverage the productive software stack and hardware board support package that HLS provides but allow for fine-grained optimization using RTL components. In this work, we introduce a new software-hardware co-design framework that integrates OpenCL/OpenACC with RTL code enabling direct execution on FPGAs as well as full emulation with a high-speed simulator to reduce the development time. Blaise-Pascal Tine, Seyong Lee, Jeffrey S. Vetter, Hyesoon Kim |
FPGA | 2 |
| 2020 | CCAMP: an integrated translation and optimization framework for OpenACC and OpenMPabstractHeterogeneous computing and exploration into specialized accelerators are inevitable in current and future supercomputers. Although this diversity of devices is promising for performance, the array of architectures presents programming challenges. High-level programming strategies have emerged to face these challenges, such as the OpenMP offloading model and OpenACC. However, the varying levels of support for these standards within vendor-specific and open-source tools, as well as the lack of performance portability across devices, have prevented the standards from achieving their goals. To address these shortcomings, we present CCAMP, an OpenMP and OpenACC interoperable framework. CCAMP provides two primary facilities: language translation between the two standards and device-specific directive optimization within each standard. We show that by using the CCAMP framework, programmers can easily transplant non-portable code into new ecosystems for new architectures. Additionally, by using CCAMP's device-specific directive optimizations, users can achieve optimized performance across architectures using a single source code. Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
SC | 2 |
| 2019 | Performance portability study for massively parallel computational fluid dynamics application on scalable heterogeneous architectures
Seyong Lee, John Gounley, Amanda Randles, Jeffrey S. Vetter |
J. Parallel Distributed Comput. | 1 |
| 2018 | Tuyere: enabling scalable memory workloads for system explorationabstractMemory technologies are under active development. Meanwhile, workloads on contemporary computing systems are increasing rapidly in size and diversity. Such dynamics in hardware and software further widen the gap between memory system design and performance evaluation. In this work, we propose a data-centric abstraction of high-performance computing applications for fast exploration of new memory technologies. We also provide a framework that uses a formal modeling language to describe the abstraction, automatically translates abstractions into memory traffic, and directly interfaces with cycle-accurate simulators. We evaluated the framework using 20 workloads and validated the memory traffic profile, the simulation results, and the relative memory changes of four memory technologies. Our results show that the data-centric abstraction can accurately capture application behavior adaptable to different input problems and can expedite system exploration. Ivy Bo Peng, Jeffrey S. Vetter, Shirley V. Moore, Seyong Lee |
HPDC | 4 |
| 2018 | Directive-Based, High-Level Programming and Optimizations for High-Performance Computing with FPGAsabstractReconfigurable architectures like Field Programmable Gate Arrays (FPGAs) have been used for accelerating computations from several domains because of their unique combination of flexibility, performance, and power efficiency. However, FPGAs have not been widely used for high-performance computing, primarily because of their programming complexity and difficulties in optimizing performance. In this paper, we present a directive-based, high-level optimization framework for high-performance computing with FPGAs, built on top of an OpenACC-to-FPGA translation framework called OpenARC. We propose directive extensions and corresponding compile-time optimization techniques to enable the compiler to generate more efficient FPGA hardware configuration files. Empirical evaluation of the proposed framework on an Intel Stratix V with five OpenACC benchmarks from various application domains shows that FPGA-specific optimizations can lead to significant increases in performance across all tested applications. We also demonstrate that applying these high-level directive-based optimizations can allow OpenACC applications to perform similarly to lower-level OpenCL applications with hand-written FPGA-specific optimizations, and offer runtime and power performance benefits compared to CPUs and GPUs. Jacob Lambert 0002, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
ICS | 2 |
| 2018 | GPU Data Access on Complex Geometries for D3Q19 Lattice Boltzmann MethodabstractGPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When LBM is advanced with GPUs on complex computational domains, geometric data is typically accessed indirectly, and lattice data is typically accessed lexicographically in the Structure of Array (SoA) layout. Although there are a variety of existing access patterns beyond the typical choices, no study has yet examined the relative efficacy between them. Here, we compare a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct addressing is the superior addressing scheme for the majority of cases examined: Semi-direct addressing increases computational speed and often reduces memory consumption. For lattice layout, we find that the Collected Structure of Arrays (CSoA) layout outperforms the SoA layout. When compared to state-of-the-art practices, our recommended addressing modifications lead to performance gains between 10-40% across different complex geometries, fluid volume fractions, and resolutions. The modifications also lead to a decrease in memory consumption by as much as 17%. Having discovered these improvements, we examine a highly resolved arterial geometry on a leadership class system. On this system we present the first near-optimal strong results for LBM with arterial geometries run on GPUs. We also demonstrate that the above recommendations remain valid for large scale, many device simulations, which leads to an increased computational speed and average memory usage reductions. To understand these observations, we employ performance modeling which reveals that semi-direct methods outperform indirect methods due to a reduced number of total loads/stores in memory, and that CSoA outperforms SoA and bundling due to improved caching behavior. Gregory Herschlag, Seyong Lee, Jeffrey S. Vetter, Amanda Randles |
IPDPS | 2 |
| 2018 | Highly Efficient Compensation-Based Parallelism for Wavefront Loops on GPUsabstractWavefront loops are widely used in many scientific applications, e.g., partial differential equation (PDE) solvers and sequence alignment tools. However, due to the data dependencies in wavefront loops, it is challenging to fully utilize the abundant compute units of GPUs and to reuse data through their memory hierarchy. Existing solutions can only optimize for these factors to a limited extent. For example, tiling-based methods optimize memory access but may result in load imbalance; while compensation-based methods, which change the original order of computation to expose more parallelism and then compensate for it, suffer from both global synchronization overhead and limited generality. In this paper, we first prove under which circumstances that breaking data dependencies and properly changing the sequence of computation operators in our compensation-based method does not affect the correctness of results. Based on this analysis, we design a highly efficient compensation-based parallelism on GPUs. Our method provides weighted scan-based GPU kernels to optimize the computation and combines with the tiling method to optimize memory access and synchronization. The performance results on the NVIDIA K80 and P100 GPU platforms demonstrate that our method can achieve significant improvements for four types of real-world application kernels over the state-of-the-art research. Kaixi Hou, Hao Wang 0002, Wu-chun Feng, Jeffrey S. Vetter, Seyong Lee |
IPDPS | 5 |
| 2018 | Juggler: a dependence-aware task-based execution framework for GPUsabstractScientific applications with single instruction, multiple data (SIMD) computations show considerable performance improvements when run on today's graphics processing units (GPUs). However, the existence of data dependences across thread blocks may significantly impact the speedup by requiring global synchronization across multiprocessors (SMs) inside the GPU. To efficiently run applications with interblock data dependences, we need fine-granular task-based execution models that will treat SMs inside a GPU as stand-alone parallel processing units. Such a scheme will enable faster execution by utilizing all internal computation elements inside the GPU and eliminating unnecessary waits during device-wide global barriers. Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Laxmi N. Bhuyan |
PPoPP | 2 |
| 2018 | DRAGON: breaking GPU memory capacity limits with direct NVM access
Pak Markthub, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Satoshi Matsuoka |
SC | 3 |
| 2018 | The OpenACC data model: Preliminary study on its major challenges and implementations
Michael Wolfe, Seyong Lee, Xiaonan Tian, Rengan Xu, Barbara M. Chapman, Sunita Chandrasekaran |
Parallel Comput. | 2 |
| 2017 | Language-Based Optimizations for Persistence on Nonvolatile Main Memory SystemsabstractSubstantial advances in nonvolatile memory (NVM) technologies have motivated wide-spread integration of NVM into mobile, enterprise, and HPC systems. Recently, considerable research has focused on architectural integration of NVM and respective programming systems, exploiting NVM's trait of persistence correctly and efficiently. In this regard, we design several novel language-based optimization techniques for programming NVM and demonstrate them as an extension of our NVL-C system. Specifically, we focus on optimizing the performance of atomic updates to complex data structures residing in NVM. We build on two variants of automatic undo logging: canonical undo logging, and shadow updates. We show these techniques can be implemented transparently and efficiently, using dynamic selection and other logging optimizations. Our empirical results on several applications gathered on an NVM testbed illustrate that our cost-model-based dynamic selection technique can accurately choose the best logging variant across different NVM modes and input sizes. In comparison to statically choosing canonical undo logging, this improvement reduces execution time to as little as 53% for block-addressable NVM and 73% for emulated byte-addressable NVM on a Fusion-io ioScale device. Joel E. Denny, Seyong Lee, Jeffrey S. Vetter |
IPDPS | 2 |
| 2017 | Design and Implementation of Papyrus: Parallel Aggregate Persistent StorageabstractA surprising development in recently announced HPC platforms is the addition of, sometimes massive amounts of, persistent (nonvolatile) memory (NVM) in order to increase memory capacity and compensate for plateauing I/O capabilities. However, there are no portable and scalable programming interfaces using aggregate NVM effectively. This paper introduces Papyrus: a new software system built to exploit emerging capability of NVM in HPC architectures. Papyrus (or Parallel Aggregate Persistent -YRU- Storage) is a novel programming system that provides features for scalable, aggregate, persistent memory in an extreme-scale system for typical HPC usage scenarios. Papyrus mainly consists of Papyrus Virtual File System (VFS) and Papyrus Template Container Library (TCL). Papyrus VFS provides a uniform aggregate NVM storage image across diverse NVM architectures. It enables Papyrus TCL to provide a portable and scalable high-level container programming interface whose data elements are distributed across multiple NVM nodes without requiring the user to handle complex communication, synchronization, replication, and consistency model. We evaluate Papyrus on two HPC systems, including UTK Beacon and NERSC Cori, using real NVM storage devices. Kittisak Sajjapongse, Seyong Lee, Jeffrey S. Vetter |
IPDPS | 3 |
| 2017 | PapyrusKV: a high-performance parallel key-value store for distributed NVM architecturesabstractThis paper introduces PapyrusKV, a parallel embedded key-value store (KVS) for distributed high-performance computing (HPC) architectures that offer potentially massive pools of nonvolatile memory (NVM). PapyrusKV stores keys with their values in arbitrary byte arrays across multiple NVMs in a distributed system. PapyrusKV provides standard KVS operations such as put, get, and delete. More importantly, PapyrusKV provides advanced features for HPC such as dynamic consistency control, zero-copy workflow, and asynchronous checkpoint/restart. Beyond filesystems, PapyrusKV provides HPC programmers with a high-level interface to exploit distributed NVM in the system, and it transparently organizes data to achieve high performance. Also, it allows HPC applications to specialize PapyrusKV to meet their specific requirements. We empirically evaluate PapyrusKV on three HPC systems with real NVM devices: OLCF's Summitdev, TACC's Stampede, and NERSC's Cori. Our results show that PapyrusKV can offer high performance, scalability, and portability across these various distributed NVM architectures. Seyong Lee, Jeffrey S. Vetter |
SC | 2 |
| 2016 | NVL-C: Static Analysis Techniques for Efficient, Correct Programming of Non-Volatile Main Memory SystemsabstractComputer architecture experts expect that non-volatile memory (NVM) hierarchies will play a more significant role in future systems including mobile, enterprise, and HPC architectures. With this expectation in mind, we present NVL-C: a novel programming system that facilitates the efficient and correct programming of NVM main memory systems. The NVL-C programming abstraction extends C with a small set of intuitive language features that target NVM main memory, and can be combined directly with traditional C memory model features for DRAM. We have designed these new features to enable compiler analyses and run-time checks that can improve performance and guard against a number of subtle programming errors, which, when left uncorrected, can corrupt NVM-stored data. Moreover, to enable recovery of data across application or system failures, these NVL-C features include a flexible directive for specifying NVM transactions. So that our implementation might be extended to other compiler front ends and languages, the majority of our compiler analyses are implemented in an extended version of LLVM's intermediate representation (LLVM IR). We evaluate NVL-C on a number of applications to show its flexibility, performance, and correctness. Joel E. Denny, Seyong Lee, Jeffrey S. Vetter |
HPDC | 2 |
| 2016 | IMPACC: A Tightly Integrated MPI+OpenACC Framework Exploiting Shared Memory ParallelismabstractWe propose IMPACC, an MPI+OpenACC framework for heterogeneous accelerator clusters. IMPACC tightly integrates MPI and OpenACC, while exploiting the shared memory parallelism in the target system. IMPACC dynamically adapts the input MPI+OpenACC applications on the target heterogeneous accelerator clusters to fully exploit target system-specific features. IMPACC provides the programmers with the unified virtual address space, automatic NUMA-friendly task-device mapping, efficient integrated communication routines, seamless streamlining of asynchronous executions, and transparent memory sharing. We have implemented IMPACC and evaluated its performance using three heterogeneous accelerator systems, including Titan supercomputer. Results show that IMPACC can achieve easier programming, higher performance, and better scalability than the current MPI+OpenACC model. Seyong Lee, Jeffrey S. Vetter |
HPDC | 2 |
| 2016 | OpenACC to FPGA: A Framework for Directive-Based High-Performance Reconfigurable ComputingabstractThis paper presents a directive-based, high-level programming framework for high-performance reconfigurable computing. It takes a standard, portable OpenACC C program as input and generates a hardware configuration file for execution on FPGAs. We implemented this prototype system using our open-source OpenARC compiler, it performs source-to-source translation and optimization of the input OpenACC program into an OpenCL code, which is further compiled into a FPGA program by the backend Altera Offline OpenCL compiler. Internally, the design of OpenARC uses a high-level intermediate representation that separates concerns of program representation from underlying architectures, which facilitates portability of OpenARC. In fact, this design allowed us to create the OpenACC-to-FPGA translation framework with minimal extensions to our existing system. In addition, we show that our proposed FPGA-specific compiler optimizations and novel OpenACC pragma extensions assist the compiler in generating more efficient FPGA hardware configuration files. Our empirical evaluation on an Altera Stratix V FPGA with eight OpenACC benchmarks demonstrate the benefits of our strategy. To demonstrate the portability of OpenARC, we show results for the same benchmarks executing on other heterogeneous platforms, including NVIDIA GPUs, AMD GPUs, and Intel Xeon Phis. This initial evidence helps support the goal of using a directive-based, high-level programming strategy for performance portability across heterogeneous HPC architectures. Seyong Lee, Jeffrey S. Vetter |
IPDPS | 1 |
| 2015 | Programmer-Guided Reliability for Extreme-Scale ApplicationsabstractWe present "programmer-guided reliability" (PGR) as a systematic conceptual approach to address the expected rise in soft errors in coming extreme-scale systems at the application level. The approach involves instrumentation of the application with code to detect data corruption errors. The location and nature of these error detectors are at the discretion of the programmer, who uses their knowledge and experience with the problem domain, the application, the solution algorithms, etc., to determine the most vulnerable areas of the code and the most appropriate ways to detect data corruption. To illustrate the approach, we provide examples of error detectors from four different benchmark-scale applications. We also describe a simple control framework that allows for runtime configuration of the error detectors without recompilation of the application, as well as dynamic reconfiguration during the execution of the application. Finally, we discuss a number of future directions building on the basic PGR approach, including the incorporation of some general error detectors into the programming environment in order to make them more easily usable by the programmer. David E. Bernholdt, Wael R. Elwasif, Christos Kartsaklis, Seyong Lee, Tiffany M. Mintz |
CLUSTER | 4 |
| 2015 | COMPASS: A Framework for Automated Performance Modeling and PredictionabstractFlexible, accurate performance predictions offer numerous benefits such as gaining insight into and optimizing applications and architectures. However, the development and evaluation of such performance predictions has been a major research challenge, due to the architectural complexities. To address this challenge, we have designed and implemented a prototype system, named COMPASS, for automated performance model generation and prediction. COMPASS generates a structured performance model from the target application's source code using automated static analysis, and then, it evaluates this model using various performance prediction techniques. As we demonstrate on several applications, the results of these predictions can be used for a variety of purposes, such as design space exploration, identifying performance tradeoffs for applications, and understanding sensitivities of important parameters. COMPASS can generate these predictions across several types of applications from traditional, sequential CPU applications to GPU-based, heterogeneous, parallel applications. Our empirical evaluation demonstrates a maximum overhead of 4%, flexibility to generate models for 9 applications, speed, ease of creation, and very low relative errors across a diverse set of architectures. Seyong Lee, Jeremy S. Meredith, Jeffrey S. Vetter |
ICS | 1 |
| 2015 | An OpenACC-based unified programming model for multi-accelerator systemsabstractThis paper proposes a novel SPMD programming model of OpenACC. Our model integrates the different granularities of parallelism from vector-level parallelism to node-level parallelism into a single, unified model based on OpenACC. It allows programmers to write programs for multiple accelerators using a uniform programming model whether they are in shared or distributed memory systems. We implement a prototype of our model and evaluate its performance with a GPU-based supercomputer using three benchmark applications. Seyong Lee, Jeffrey S. Vetter |
PPoPP | 2 |
| 2014 | OpenARC: open accelerator research compiler for directive-based, efficient heterogeneous computingabstractThis paper presents Open Accelerator Research Compiler (OpenARC): an open-source framework that supports the full feature set of OpenACC V1.0 and performs source-to-source transformations, targeting heterogeneous devices, such as NVIDIA GPUs. Combined with its high-level, extensible Intermediate Representation (IR) and rich semantic annotations, OpenARC serves as a powerful research vehicle for prototyping optimization, source-to-source transformations, and instrumentation for debugging, performance analysis, and autotuning. In fact, OpenARC is equipped with various capabilities for advanced analyses and transformations, as well as built-in performance and debugging tools. We explain the overall design and implementation of OpenARC, and we present key analysis techniques necessary to efficiently port OpenACC applications. Porting various OpenACC applications to CUDA GPUs using OpenARC demonstrates that OpenARC performs similarly to a commercial compiler, while serving as a general research framework. Seyong Lee, Jeffrey S. Vetter |
HPDC | 1 |
| 2014 | Interactive Program Debugging and Optimization for Directive-Based, Efficient GPU ComputingabstractDirective-based GPU programming models are gaining momentum, since they transparently relieve programmers from dealing with complexity of low-level GPU programming, which often reflects the underlying architecture. However, too much abstraction in directive models puts a significant burden on programmers for debugging applications and tuning performance. In this paper, we propose a directive-based, interactive program debugging and optimization system. This system enables intuitive and synergistic interaction among programmers, compilers, and runtimes for more productive and efficient GPU computing. We have designed and implemented a series of prototype tools within our new open source compiler framework, called Open Accelerator Research Compiler (Open ARC), Open ARC supports the full feature set of Opencast V1.0. Our evaluation on twelve Open ACC benchmarks demonstrates that our prototype debugging and optimization system can detect a variety of translation errors. Additionally, the optimization provided by our prototype minimizes memory transfers, when compared to a fully manual memory management scheme. Seyong Lee, Dong Li 0001, Jeffrey S. Vetter |
IPDPS | 1 |
| 2013 | MapReduce with communication overlap (MaRCO)
Faraz Ahmad, Seyong Lee, Mithuna Thottethodi, T. N. Vijaykumar |
J. Parallel Distributed Comput. | 2 |
| 2012 | Early evaluation of directive-based GPU programming models for productive exascale computingabstractGraphics Processing Unit (GPU)-based parallel computer architectures have shown increased popularity as a building block for high performance computing, and possibly for future Exascale computing. However, their programming complexity remains as a major hurdle for their widespread adoption. To provide better abstractions for programming GPU architectures, researchers and vendors have proposed several directive-based GPU programming models. These directive-based models provide different levels of abstraction, and required different levels of programming effort to port and optimize applications. Understanding these differences among these new models provides valuable insights on their applicability and performance potential. In this paper, we evaluate existing directive-based models by porting thirteen application kernels from various scientific domains to use CUDA GPUs, which, in turn, allows us to identify important issues in the functionality, scalability, tunability, and debuggability of the existing models. Our evaluation shows that directive-based models can achieve reasonable performance, compared to hand-written GPU codes. Seyong Lee, Jeffrey S. Vetter |
SC | 1 |
| 2010 | OpenMPC: Extended OpenMP Programming and Tuning for GPUsabstractGeneral-Purpose Graphics Processing Units (GPGPUs) are promising parallel platforms for high performance computing. The CUDA (Compute Unified Device Architecture) programming model provides improved programmability for general computing on GPGPUs. However, its unique execution model and memory model still pose significant challenges for developers of efficient GPGPU code. This paper proposes a new programming interface, called OpenMPC, which builds on OpenMP to provide an abstraction of the complex CUDA programming model and offers high-level controls of the involved parameters and optimizations. We have developed a fully automatic compilation and user-assisted tuning system supporting OpenMPC. In addition to a range of compiler transformations and optimizations, the system includes tuning capabilities for generating, pruning, and navigating the search space of compilation variants. Our results demonstrate that OpenMPC offers both programmability and tunability. Our system achieves 88% of the performance of the hand-coded CUDA programs. Seyong Lee, Rudolf Eigenmann |
SC | 1 |
| 2009 | OpenMP to GPGPU: a compiler framework for automatic translation and optimizationabstractGPGPUs have recently emerged as powerful vehicles for general-purpose high-performance computing. Although a new Compute Unified Device Architecture (CUDA) programming model from NVIDIA offers improved programmability for general computing, programming GPGPUs is still complex and error-prone. This paper presents a compiler framework for automatic source-to-source translation of standard OpenMP applications into CUDA-based GPGPU applications. The goal of this translation is to further improve programmability and make existing OpenMP applications amenable to execution on GPGPUs. In this paper, we have identified several key transformation techniques, which enable efficient GPU global memory access, to achieve high performance. Experimental results from two important kernels (JACOBI and SPMUL) and two NAS OpenMP Parallel Benchmarks (EP and CG) show that the described translator and compile-time optimizations work well on both regular and irregular applications, leading to performance improvements of up to 50X over the unoptimized translation (up to 328X over serial). Seyong Lee, Seung-Jai Min, Rudolf Eigenmann |
PPoPP | 1 |
| 2008 | Adaptive runtime tuning of parallel sparse matrix-vector multiplication on distributed memory systemsabstractSparse matrix-vector (SpMV) multiplication is a widely used kernel in scientific applications. In these applications, the SpMV multiplication is usually deeply nested within multiple loops and thus executed a large number of times. We have observed that there can be significant performance variability, due to irregular memory access patterns. Static performance optimizations are difficult because the patterns may be known only at runtime. In this paper, we propose adaptive runtime tuning mechanisms to improve the parallel performance on distributed memory systems. Our adaptive iteration-to-process mapping mechanism balances computational load at runtime with negligible overhead (1% on average), and our runtime communication selection algorithm searches for the best communication method for a given data distribution and mapping. Actual runs on 26 real matrices show that our runtime tuning system reduces execution time up to 68.8% (30.9% on average) over a base block-distributed parallel algorithm on distributed systems with 32 nodes. Seyong Lee, Rudolf Eigenmann |
ICS | 1 |
| 2008 | Adaptive tuning in a dynamically changing resource environmentabstractWe present preliminary results of a project to create a tuning system that adaptively optimizes programs to the underlying execution platform. We will show initial results from two related efforts, (i) Our tuning system can efficiently select the best combination of compiler options, when translating programs to a target system, (ii) By tuning irregular applications that operate on sparse matrices, our system is able to achieve substantial performance improvements on cluster platforms. This project is part of a larger effort that aims at creating a global information sharing system, where resources, such as software applications, computer platforms, and information can be shared, discovered, and adapted to local needs. Seyong Lee, Rudolf Eigenmann |
IPDPS | 1 |
| 2008 | Efficient content search in iShare, a P2P based Internet-sharing systemabstractThis paper presents an efficient content search system, which is applied to iShare, a distributed peer-to-peer(P2P) Internet-sharing system. iShare facilitates the sharing of diverse resources located in different administrative domains over the Internet. For efficient resource management, iShare organizes resources into a hierarchical name space, which is distributed over the underlying structured P2P network. However, iShare's search capability has a fundamental limit inherited from the underlying structured P2P system's search capability. Most existing structured P2P systems do not support content searches. There exists some research that provides content search functionality, but the approaches do not scale well and incur substantial overheads on data updates. To address these issues, we propose an efficient hierarchical-summary system, which enables an efficient content search and semantic ranking capability over traditional structured P2P systems. Our system uses a hierarchical name space to implement a summary hierarchy on top of existing structured P2P overlay networks, and uses a Bloom Filter as a summary structure to reduce space and maintenance overhead. We implemented the proposed system in iShare, and the results show that our search system finds all relevant results regardless of summary scale and the search latency increases very slowly as the network grows. Seyong Lee, Xiaojuan Ren, Rudolf Eigenmann |
IPDPS | 1 |
| 2007 | Prediction of Resource Availability in Fine-Grained Cycle Sharing Systems Empirical Evaluation
Xiaojuan Ren, Seyong Lee, Rudolf Eigenmann, Saurabh Bagchi |
J. Grid Comput. | 2 |
| 2006 | Resource Availability Prediction in Fine-Grained Cycle Sharing SystemsabstractFine-grained cycle sharing (FGCS) systems aim at utilizing the large amount of computational resources available on the Internet. In FGCS, host computers allow guest jobs to utilize the CPU cycles if the jobs do not significantly impact the local users of a host. A characteristic of such resources is that they are generally provided voluntarily and their availability fluctuates highly. Guest jobs may fail because of unexpected resource unavailability. To provide fault tolerance to guest jobs without adding significant computational overhead, it requires to predict future resource availability. This paper presents a method for resource availability prediction in FGCS systems. It applies a semi-Markov Process and is based on a novel resource availability model, combining generic hardware-software failures with domain-specific resource behavior in FGCS. We describe the prediction framework and its implementation in a production FGCS system named iShare. Through the experiments on an iShare testbed, we demonstrate that the prediction achieves accuracy above 86% on average and outperforms linear time series models, while the computational cost is negligible. Our experimental results also show that the prediction is robust in the presence of irregular resource unavailability Xiaojuan Ren, Seyong Lee, Rudolf Eigenmann, Saurabh Bagchi |
HPDC | 2 |