Ivy Bo Peng

dblp:165/0083 · also Ivy Peng · DBLP profile ↗
← Back
43ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0003-4158-3583ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 8 first-author · 15 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
Ruimin Shi, Maya B. Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Bo Peng
Euro-Par (1)5
2026 Communication Offloading on SmartNIC DPUs: A Quantitative Approach
Jacob Wahlgren, Andong Hu, Roger A. Pearce, Maya B. Gokhale, Ivy Bo Peng
Euro-Par (1)5
2026 High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
Ruimin Shi, Gabin Schieffer, Pei-Hung Lin, Maya B. Gokhale, Andreas Herten, Ivy Bo Peng
IPDPS6
2026 A hybrid quantum-classical particle-in-cell method for plasma simulations
abstract
We present a hybrid quantum-classical electrostatic Particle-in-Cell (PIC) method, where the electrostatic field Poisson solver is implemented on a quantum computer simulator using a hybrid classical-quantum Neural Network (HNN) using data-driven and physics-informed learning approaches. The HNN is trained on classical PIC simulation results and executed via a PennyLane quantum simulator. The remaining computational steps, including particle motion and field interpolation, are performed on a classical system. To evaluate the accuracy and computational cost of this hybrid approach, we test the hybrid quantum-classical electrostatic PIC against the two-stream instability, a standard benchmark in plasma physics. Our results show that the quantum Poisson solver achieves comparable accuracy to classical methods. It also provides insights into the feasibility of using quantum computing and HNNs for plasma simulations. We also discuss the computational overhead associated with current quantum computer simulators, showing the challenges and potential advantages of hybrid quantum-classical numerical methods.
Pratibha Raghupati Hegde, Paolo Marcandelli, Yuanchun He, Luca Pennati, Jeremy J. Williams, Ivy Bo Peng, Stefano Markidis
Future Gener. Comput. Syst.6
2026 Tightly-integrated quantum-classical computing using the QHDL hardware description language
abstract
We present the design, development, and application of QHDL, a quantum hardware description language specifically designed for tightly-coupled quantum–classical computing systems. Together with the language design principles, we describe the QHDL compiler, debugger, and co-simulation infrastructure. We showcase the benefits of using a quantum–classical integrated approach in four use cases, requiring close quantum–classical device interaction: Bell’s pair circuit, dynamic delay, Quantum Fourier Transform (QFT), and teleportation. To interface with QHDL, we propose to use synchronous techniques that are commonplace in digital hardware design. We illustrate examples of modelling both loosely-coupled and tightly-coupled quantum circuits that use so-called measurement-in-the-middle by utilizing these techniques in QHDL. For clock-cycle accurate implementations, we propose implementing such modules as programmable hardware blocks using Register-Transfer Level (RTL) or gate-level approaches. These approaches provide the highest coupling performance and are feasible to be implemented in state-of-the-art control systems.
Gilbert Netzer, Pratibha Raghupati Hegde, Ivy Bo Peng, Stefano Markidis
Future Gener. Comput. Syst.3
2025 Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El Capitan
abstract
We present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures.
Ian Lumsden, Stefano Markidis, Andong Hu, Ivy Bo Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Thomas Scogland, Hariharan Devarajan, Bronis R. de Supinski, Gian Luca Delzanno, Michela Taufer
eScience4
2025 ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
Daniel Araújo de Medeiros, Jeremy J. Williams, Jacob Wahlgren, Leonardo Saud Maia Leite, Ivy Bo Peng
Euro-Par (1)5
2025 ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Ruimin Shi, Gabin Schieffer, Maya B. Gokhale, Pei-Hung Lin, Hiren D. Patel, Ivy Bo Peng
Euro-Par (2)6
2025 Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
abstract
Graphics Processing Units (GPUs) have become the standard in accelerating scientific applications on heterogeneous systems. However, as GPUs are getting faster, one potential performance bottleneck with GPU-accelerated applications is the overhead from launching several fine-grained kernels. CUDA Graph addresses these performance challenges by enabling a graph-based execution model that captures operations as nodes and dependence as edges in a static graph. Thereby consolidating several kernel launches into one graph launch. We propose a performance optimization strategy for iteratively launched kernels. By grouping kernel launches into iteration batches and then unrolling these batches into a CUDA Graph, iterative applications can benefit from CUDA Graph for performance boosting. We analyze the performance gain and overhead from this approach by designing a skeleton application. The skeleton application also serves as a generalized example of converting an iterative solver to CUDA Graph, and for deriving a performance model. Using the skeleton application, we show that when unrolling iteration batches for a given platform, there is an optimal size of the iteration batch, which is independent of workload, balancing the extra overhead from graph creation with the performance gain of the graph execution. Depending on workload, we show that the optimal iteration batch size gives more than $1.4 \times$ speed-up in the skeleton application. Furthermore, we show that similar speed-up can be gained in Hotspot and Hotspot3D from the Rodinia benchmark suite and a Finite-Difference Time-Domain (FDTD) Maxwell solver.
Jonah Ekelund, Stefano Markidis, Ivy Bo Peng
PDP3
2025 Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
abstract
Quantum computer simulators are an indispensable tool for prototyping quantum algorithms and verifying the functioning of existing quantum computer hardware. The current largest quantum computers feature more than one thousand qubits, challenging their classical simulators. Statevector quantum simulators are challenged by the exponential increase of representable quantum states with respect to the number of qubits, making more than fifty qubits practically unfeasible. A more appealing approach for simulating quantum computers is adopting the tensor network approach, whose memory requirements fundamentally depend on the level of entanglement in the quantum circuit, and allows simulating the current largest quantum computers. This work investigates and evaluates the CUDA-Q tensor network simulators on an Nvidia Grace Hopper system, particularly the Matrix Product State (MPS) formulation. We compare the performance of the CUDA-Q state vector implementation and validate the correctness of MPS simulations. Our results highlight that tensor network-based methods provide a significant opportunity to simulate large-qubit circuits, albeit approximately. We also show that current GPUaccelerated computation cannot fully utilize GPU efficiently in the case of MPS simulations.
Gabin Schieffer, Stefano Markidis, Ivy Bo Peng
PDP3
2024 Beyond the Buzz: Strategic Paths for Enabling Useful NISQ Applications
abstract
There is much debate on whether quantum computing on current NISQ devices, consisting of noisy hundred qubits and requiring a non-negligible usage of classical computing as part of the algorithms, has utility and will ever offer advantages for scientific and industrial applications with respect to traditional computing. In this position paper, we argue that while real-world NISQ quantum applications have yet to surpass their classical counterparts, strategic approaches can be used to facilitate advancements in both industrial and scientific applications. We have identified three key strategies to guide NISQ computing towards practical and useful implementations. Firstly, prioritizing the identification of a "killer app" is a key point. An application demonstrating the distinctive capabilities of NISQ devices can catalyze broader development. We suggest focusing on applications that are inherently quantum, e.g., pointing towards quantum chemistry and material science as promising domains. These fields hold the potential to exhibit benefits, setting benchmarks for other applications to follow. Secondly, integrating AI and deep-learning methods into NISQ computing is a promising approach. Examples such as quantum Physics-Informed Neural Networks and Differentiable Quantum Circuits (DQC) demonstrate the synergy between quantum computing and AI. Lastly, recognizing the interdisciplinary nature of NISQ computing, we advocate for a co-design approach. Achieving synergy between classical and quantum computing necessitates an effort in co-designing quantum applications, algorithms, and programming environments, and the integration of HPC with quantum hardware. The interoperability of these components is crucial for enabling the full potential of NISQ computing. In conclusion, through the usage of these three approaches, we argue that NISQ computing can surpass current limitations and evolve into a valuable tool for scientific and industrial applications. This requires an approach that integrates domain-specific killer apps, harnesses the power of quantum-enhanced AI, and embraces a collaborative co-design methodology.
Pratibha Raghupati Hegde, Oleksandr Kyriienko, Hermanni Heimonen, Panagiotis Tolias, Gilbert Netzer, Panagiotis Kl. Barkoutsos, Ricardo Vinuesa, Ivy Bo Peng, Stefano Markidis
CF8
2024 Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
abstract
Memory management across discrete CPU and GPU physical memory is traditionally achieved through explicit GPU allocations and data copy or unified virtual memory. The Grace Hopper Superchip, for the first time, supports an integrated CPU-GPU system page table, hardware-level addressing of system allocated memory, and cache-coherent NVLink-C2C interconnect, bringing an alternative solution for enabling a Unified Memory system. In this work, we provide the first in-depth study of the system memory management on the Grace Hopper Superchip, in both in-memory and memory oversubscription scenarios. We provide a suite of six representative applications, including the Qiskit quantum computing simulator, using system memory and managed memory. Using our memory utilization profiler and hardware counters, we quantify and characterize the impact of the integrated CPU-GPU system page table on GPU applications. Our study focuses on first-touch policy, page table entry initialization, page sizes, and page migration. We identify practical optimization strategies for different access patterns. Our results show that as a new solution for unified memory, the system-allocated memory can benefit most use cases with minimal porting efforts.
Gabin Schieffer, Jacob Wahlgren, Jie Ren 0015, Jennifer Faj, Ivy Bo Peng
ICPP5
2024 On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability
abstract
Matrix multiplication is a core computational part of deep learning and scientific workloads. The emergence of Matrix Cores in high-end AMD GPUs, a building block of Exascale computers, opens new opportunities for optimizing the performance and power efficiency of compute-intensive applications. This work provides a timely, comprehensive characterization of the novel Matrix Cores in AMD GPUs. We develop low-level micro-benchmarks for leveraging Matrix Cores at different levels of parallelism, achieving up to 350, 88, and 69 TFLOPS for mixed, float, and double precision on one GPU. Using results obtained from the micro-benchmarks, we provide a performance model of Matrix Cores that can guide application developers in performance tuning. We also provide the first quantitative study and modeling of the power efficiency of Matrix Cores at different floating-point data types. Finally, we evaluate the high- level programmability of Matrix Cores through the rocBLAS library in a wide range of matrix sizes from 16 to 64K. Our results indicate that application developers can transparently leverage Matrix Cores to deliver more than 92% peak computing throughput by properly selecting data types and interfaces.
Gabin Schieffer, Daniel Araújo de Medeiros, Jennifer Faj, Aniruddha Marathe, Ivy Bo Peng
ISPASS5
2024 Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
abstract
Disaggregated memory breaks the boundary of monolithic servers to enable memory provisioning on demand. Using network-attached memory to provide memory expansion for memory-intensive applications on compute nodes can improve the overall memory utilization on a cluster and reduce the total cost of ownership. However, current software solutions for leveraging network-attached memory must consume resources on the compute node for memory management tasks. Emerging off-path smartNICs provide general-purpose programmability at low-cost low-power cores. This work provides a general architecture design that enables network-attached memory and offloading tasks onto off-path programmable SmartNIC. We provide a prototype implementation called SODA on Nvidia BlueField DPU. SODA adapts communication paths and data transfer alternatives, pipelines data movement stages, and enables customizable data caching and prefetching optimizations. We evaluate SODA in five representative graph applications on real-world graphs. Our results show that SODA can achieve up to 7.9x speedup compared to node-local SSD and reduce network traffic by 42 % compared to disaggregated memory without SmartNIC offloading at similar or better performance.
Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Roger A. Pearce, Ivy Bo Peng
SBAC-PAD5
2023 Quantum Computer Simulations at Warp Speed: Assessing the Impact of GPU Acceleration: A Case Study with IBM Qiskit Aer, Nvidia Thrust & cuQuantum
abstract
Quantum computer simulators are crucial for the development of quantum computing. This work investigates GPU and multi-GPU systems' suitability and performance impact on a widely used simulation tool – the state vector simulator Qiskit Aer. In particular, we evaluate the performance of both Qiskit's default Nvidia Thrust backend and the recent Nvidia cuQuantum backend on Nvidia A100 GPUs. We provide a benchmark suite of representative quantum applications for characterization. For simulations with a large number of qubits, the two GPU backends can provide up to 14× speedup over the CPU backend, with Nvidia cuQuantum providing a further 1.5–3× speedup over the default Thrust backend. Our evaluation on a single GPU identifies the most important functions in Nvidia Thrust and cuQuantum for different quantum applications and their compute and memory bottlenecks. We also evaluate the gate fusion and cache-blocking optimizations on different quantum applications. Finally, we evaluate large-number qubit quantum applications on multi-GPU and identify data movement between host and GPU as the limiting factor for the performance.
Jennifer Faj, Ivy Bo Peng, Jacob Wahlgren, Stefano Markidis
e-Science2
2023 Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
Gabin Schieffer, Ivy Bo Peng
Euro-Par2
2023 LibCOS: Enabling Converged HPC and Cloud Data Stores with MPI
abstract
Recently, federated HPC and cloud resources are becoming increasingly strategic for providing diversified and geographically available computing resources. However, accessing data stores across HPC and cloud storage systems is challenging. Many cloud providers use object storage systems to support their clients in storing and retrieving data over the internet. One popular method is REST APIs atop the HTTP protocol, with Amazon’s S3 APIs being supported by most vendors. In contrast, HPC systems are contained within their networks and tend to use parallel file systems with POSIX-like interfaces. This work addresses the challenge of diverse data stores on HPC and cloud systems by providing native object storage support through the unified MPI I/O interface in HPC applications. In particular, we provide a prototype library called LibCOS that transparently enables MPI applications running on HPC systems to access object storage on remote cloud systems. We evaluated LibCOS on a Ceph object storage system and a traditional HPC system. In addition, we conducted performance characterization of core S3 operations that enable individual and collective MPI I/O. Our evaluation in HACC, IOR, and BigSort shows that enabling diverse data stores on HPC and Cloud storage is feasible and can be transparently achieved through the widely adopted MPI I/O. Also, we show that a native object storage system like Ceph could improve the scalability of I/O operations in parallel applications.
Daniel Araújo de Medeiros, Stefano Markidis, Ivy Bo Peng
HPC Asia3
2023 Kub: Enabling Elastic HPC Workloads on Containerized Environments
abstract
The conventional model of resource allocation in HPC systems is static. Thus, a job cannot leverage newly available resources in the system or release underutilized resources during the execution. In this paper, we present Kub, a methodology that enables elastic execution of HPC workloads on Kubernetes so that the resources allocated to a job can be dynamically scaled during the execution. One main optimization of our method is to maximize the reuse of the originally allocated resources so that the disruption to the running job can be minimized. The scaling procedure is coordinated among nodes through remote procedure calls on Kubernetes for deploying workloads in the cloud. We evaluate our approach using one synthetic benchmark and two production-level MPI-based HPC applications - GRO-MACS and CM1. Our results demonstrate that the benefits of adapting the allocated resources depend on the workload characteristics. In the tested cases, a properly chosen scaling point for increasing resources during execution achieved up to 2x speedup. Also, the overhead of checkpointing and data reshuffling significantly influences the selection of optimal scaling points and requires application-specific knowledge.
Daniel Araújo de Medeiros, Jacob Wahlgren, Gabin Schieffer, Ivy Bo Peng
SBAC-PAD4
2023 A Quantitative Approach for Adopting Disaggregated Memory in HPC Systems
abstract
Memory disaggregation has recently been adopted in data centers to improve resource utilization, motivated by cost and sustainability. Recent studies on large-scale HPC facilities have also highlighted memory underutilization. A promising and non-disruptive option for memory disaggregation is rack-scale memory pooling, where node-local memory is supplemented by shared memory pools. This work outlines the prospects and requirements for adoption and clarifies several misconceptions. We propose a quantitative method for dissecting application requirements on the memory system from the top down in three levels, moving from general, to multi-tier memory systems, and then to memory pooling. We provide a multi-level profiling tool and LBench to facilitate the quantitative approach. We evaluate a set of representative HPC workloads on an emulated platform. Our results show that prefetching activities can significantly influence memory traffic profiles. Interference in memory pooling has varied impacts on applications, depending on their access ratios to memory tiers and arithmetic intensities. Finally, in two case studies, we show the benefits of our findings at the application and system levels, achieving 50% reduction in remote access and 13% speedup in BFS, and reducing performance variation of co-located workloads in interference-aware job scheduling.
Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy Bo Peng
SC4
2022 Enabling Scalable and Extensible Memory-Mapped Datastores in Userspace
abstract
Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we presentUMap, a scalable and extensible userspace service for memory-mapping datastores. Through decoupled queue management, concurrency aware adaptation, and dynamic load balancing,UMapenables application performance to scale even at high concurrency. We evaluateUMapin data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show thatUMapas a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9${\times}$. We performed two case studies to demonstrateUMap's extensibility. First, a new datastore residing in remote memory is incorporated intoUMapas an application-specific plugin. Second, we present a persistent memory allocatorMetallbuilt atopUMapfor unified storage/memory.
Ivy Bo Peng, Maya B. Gokhale, Karim Youssef, Keita Iwabuchi, Roger A. Pearce
IEEE Trans. Parallel Distributed Syst.1
2021 ArchTM: Architecture-Aware, High Performance Transaction for Persistent Memory
Kai Wu 0006, Jie Ren 0015, Ivy Bo Peng, Dong Li 0001
FAST3
2021 Optimizing large-scale plasma simulations on persistent memory-based heterogeneous memory with effective data placement across memory hierarchy
abstract
Particle simulations of plasma are important for understanding plasma dynamics in space weather and fusion devices. However, production simulations that use billions and even trillions of computational particles require high memory capacity. In this work, we explore the latest persistent memory (PM) hardware to enable large-scale plasma simulations at unprecedented scales on a single machine. We use WarpX, an advanced plasma simulation code which is mission-critical and targets future exascale systems. We analyze the performance of WarpX on PM-based heterogeneous memory systems and propose to make the best use of memory hierarchy to avoid the impact of inferior performance of PM. We introduce a combination of static and dynamic data placement, and processor-cache prefetch mechanism for performance optimization. We develop a performance model to enable efficient data migration between PM and DRAM in the background, without reducing available bandwidth and parallelism to the application threads. We also build an analytical model to decide when to prefetch for the best use of caches. Our design achieves 66.4% performance improvement over the PM-only baseline and outperforms DRAM-cached, NUMA first-touch, and a state-of-the-art software solution by 38.8%, 45.1% and 83.3%, respectively.
Jie Ren 0015, Jiaolin Luo, Ivy Bo Peng, Kai Wu 0006, Dong Li 0001
ICS3
2021 MD-HM: memoization-based molecular dynamics simulations on big memory system
abstract
Molecular dynamics (MD) simulation is a fundamental method for modeling ensembles of particles. In this paper, we introduce a new method to improve the performance of MD by leveraging the emerging TB-scale big memory system. In particular, we trade memory capacity for computation capability to improve MD performance by the lookup table-based memoization technique. The traditional memoization technique for the MD simulation uses relatively small DRAM, bases on a suboptimal data structure, and replaces pair-wise computation, which leads to limited performance benefit in the big memory system. We introduce MD-HM, a memoization-based MD simulation framework customized for the big memory system. MD-HM partitions the simulation field into subgrids, and replaces computation in each subgrid as a whole based on a lightweight pattern-match algorithm to recognize computation in the subgrid. MD-HM uses a new two-phase LSM-tree to optimize read/write performance. Evaluating with nine MD simulations, we show that MD-HM outperforms the state-of-the-art LAMMPS simulation framework with an average speedup of 7.6x based on the Intel Optane-based big memory system.
Wenqian Dong, Jie Liu 0096, Ivy Bo Peng, Yanbao Ma, Dong Li 0001
ICS4
2020 Ribbon: High Performance Cache Line Flushing for Persistent Memory
abstract
Cache line flushing (CLF) is a fundamental building block for programming persistent memory (PM). CLF is prevalent in PM-aware workloads to ensure crash consistency. It also imposes high overhead. Extensive works have explored persistency semantics and CLF policies, but few have looked into the CLF mechanism. This work aims to improve the performance of CLF mechanism based on the performance characterization of well-established workloads on real PM hardware. We reveal that the performance of CLF is highly sensitive to the concurrency of CLF and cache line status.
Kai Wu 0006, Ivy Bo Peng, Jie Ren 0015, Dong Li 0001
PACT2
2020 ATMem: adaptive data placement in graph applications on heterogeneous memories
abstract
Active development in new memory devices, such as non-volatile memories and high-bandwidth memories, brings heterogeneous memory systems (HMS) as a promising solution for implementing large-scale memory systems with cost, area, and power limitations. Typical HMS consists of a small-capacity high-performance memory and a large-capacity low-performance memory. Data placement on such systems plays a critical role in performance optimization. Existing efforts have explored coarse-grained data placement in applications with dense data structures; however, a thorough study of applications that are based on graph data structures is still missing.
Yu Chen 0036, Ivy Bo Peng, Xu Liu 0001, Bin Ren 0002
CGO2
2020 tf-Darshan: Understanding Fine-grained I/O Performance in Machine Learning Workloads
abstract
Machine Learning applications on HPC systems have been gaining popularity in recent years. The upcoming large scale systems will offer tremendous parallelism for training through GPUs. However, another heavy aspect of Machine Learning is I/O, and this can potentially be a performance bottleneck. TensorFlow, one of the most popular Deep-Learning platforms, now offers a new profiler interface and allows instrumentation of TensorFlow operations. However, the current profiler only enables analysis at the TensorFlow platform level and does not provide system-level information. In this paper, we extend TensorFlow Profiler and introduce tf-Darshan, both a profiler and tracer, that performs instrumentation through Darshan. We use the same Darshan shared instrumentation library and implement a runtime attachment without using a system preload. We can extract Darshan profiling data structures during TensorFlow execution to enable analysis through the TensorFlow profiler. We visualize the performance results through TensorBoard, the web-based TensorFlow visualization tool. At the same time, we do not alter Darshan's existing implementation. We illustrate tf-Darshan by performing two case studies on ImageNet image and Malware classification. We show that by guiding optimization using data from tf-Darshan, we increase POSIX I/O bandwidth by up to 19% by selecting data for staging on fast tier storage. We also show that Darshan has the potential of being used as a runtime library for profiling and providing information for future optimization.
Steven W. D. Chien, Artur Podobas, Ivy Bo Peng, Stefano Markidis
CLUSTER3
2020 Demystifying the Performance of HPC Scientific Applications on NVM-based Memory Systems
abstract
The emergence of high-density byte-addressable non-volatile memory (NVM) is promising to accelerate data-and compute-intensive applications. Current NVM technologies have lower performance than DRAM and, thus, are often paired with DRAM in a heterogeneous main memory. Recently, byte-addressable NVM hardware becomes available. This work provides a timely evaluation of representative HPC applications from the "Seven Dwarfs" on NVM-based main memory. Our results quantify the effectiveness of DRAM-cached-NVM for accelerating HPC applications and enabling large problems beyond the DRAM capacity. On uncached-NVM, HPC applications exhibit three tiers of performance sensitivity, i.e., insensitive, scaled, and bottlenecked. We identify write throttling and concurrency control as the priorities in optimizing applications. We highlight that concurrency change may have a diverging effect on read and write accesses in applications. Based on these findings, we explore two optimization approaches. First, we provide a prediction model that uses datasets from a small set of configurations to estimate performance at various concurrency and data sizes to avoid exhaustive search in the configuration space. Second, we demonstrate that write-aware data placement on uncached-NVM could achieve 2x performance improvement with a 60% reduction in DRAM usage.
Ivy Bo Peng, Kai Wu 0006, Jie Ren 0015, Dong Li 0001, Maya B. Gokhale
IPDPS1
2020 sputniPIC: An Implicit Particle-in-Cell Code for Multi-GPU Systems
abstract
Large-scale simulations of plasmas are essential for advancing our understanding of fusion devices, space, and astrophysical systems. Particle-in-Cell (PIC) codes have demonstrated their success in simulating numerous plasma phenomena on HPC systems. Today, flagship supercomputers feature multiple GPUs per compute node to achieve unprecedented computing power at high power efficiency. PIC codes require new algorithm design and implementation for exploiting such accelerated platforms. In this work, we design and optimize a three-dimensional implicit PIC code, called sputniPIC, to run on a general multi-GPU compute node. We introduce a particle decomposition data layout, in contrast to domain decomposition on CPU-based implementations, to use particle batches for overlapping communication and computation on GPUs. sputniPIC also natively supports different precision representations to achieve speed up on hardware that supports reduced precision. We validate sputniPIC through the well-known GEM challenge and provide performance analysis. We test sputniPIC on three multi-GPU platforms and report a 200-800x performance improvement with respect to the sputniPIC CPU OpenMP version performance. We show that reduced precision could further improve performance by 45% to 80% on the three platforms. Because of these performance improvements, on a single node with multiple GPUs, sputniPIC enables large-scale three-dimensional PIC simulations that were only possible using clusters.
Steven W. D. Chien, Jonas Nylund, Gabriel Bengtsson, Ivy Bo Peng, Artur Podobas, Stefano Markidis
SBAC-PAD4
2020 On the Memory Underutilization: Exploring Disaggregated Memory on HPC Systems
abstract
Large-scale high-performance computing (HPC) systems consist of massive compute and memory resources tightly coupled in nodes. We perform a large-scale study of memory utilization on four production HPC clusters. Our results show that more than 90% of jobs utilize less than 15% of the node memory capacity, and for 90% of the time, memory utilization is less than 35%. Recently, disaggregated architecture is gaining traction because it can selectively scale up a resource and improve resource utilization. Based on these observations, we explore using disaggregated memory to support memory-intensive applications, while most jobs remain intact on HPC systems with reduced node memory. We designed and developed a user-space remote-memory paging library to enable applications exploring disaggregated memory on existing HPC clusters. We quantified the impact of access patterns and network connectivity in benchmarks. Our case studies of graph-processing and Monte-Carlo applications evaluated the impact of application characteristics and local memory capacity and highlighted the potential of throughput scaling on disaggregated memory.
Ivy Bo Peng, Roger A. Pearce, Maya B. Gokhale
SBAC-PAD1
2019 Analyzing the suitability of contemporary 3D-stacked PIM architectures for HPC scientific applications
abstract
Scaling off-chip bandwidth is challenging due to fundamental limitations, such as a fixed pin count and plateauing signaling rates. Recently, vendors have turned to 2.5D and 3D stacking to closely integrate system components. Interestingly, these technologies can integrate a logic layer under multiple memory dies, enabling computing capability inside a memory stack. This trend in stacking is making PIM architectures commercially viable. In this work, we investigate the suitability of offloading kernels in scientific applications onto 3D stacked PIM architectures. We evaluate several hardware constraints resulted from the stacked structure. We perform extensive simulation experiments and in-depth analysis to quantify the impact of application locality in TLBs, data caches, and memory stacks. Our results also identify design optimization areas in software and hardware for HPC scientific applications.
Ivy Bo Peng, Jeffrey S. Vetter, Shirley Moore, Joydeep Rakshit, Stefano Markidis
CF1
2019 SAGE: Percipient Storage for Exascale Data Centric Computing
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Stefano Markidis, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Dirk Pleiter, Shaun De Witt
Parallel Comput.7
2018 Understanding scale-Dependent soft-Error Behavior of Scientific Applications
abstract
Analyzing application fault behavior on large-scale systems is time-consuming and resource-demanding. Currently, researchers need to perform fault injection campaigns at full scale to understand the effects of soft errors on applications and whether these faults result in silent data corruption. Both time and resource requirements greatly limit the scope of the resilience studies that can be currently performed. In this work, we propose a methodology to model application fault behavior at large scale based on a reduced set of experiments performed at small scale. We employ machine learning techniques to accurately model application fault behavior using a set of experiments that can be executed in parallel at small scale. Our methodology drastically reduces the set and the scale of the fault injection experiments to be performed and provides a validated methodology to study application fault behavior at large scale. We show that our methodology can accurately model application fault behavior at large scale by using only small scale experiments. In some cases, we can model the fault behavior of a parallel application running on 4,096 cores with about 90% accuracy based on experiments on a single core.
Gokcen Kestor, Ivy Bo Peng, Roberto Gioiosa, Sriram Krishnamoorthy
CCGrid2
2018 The SAGE project: a storage centric approach for exascale computing: invited paper
abstract
SAGE (Percipient StorAGe for Exascale Data Centric Computing) is a European Commission funded project towards the era of Exascale computing. Its goal is to design and implement a Big Data/Extreme Computing (BDEC) capable infrastructure with associated software stack. The SAGE system follows a storage centric approach as it is capable of storing and processing large data volumes at the Exascale regime.
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Steven W. D. Chien, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Shaun De Witt, Dirk Pleiter, Stefano Markidis
CF7
2018 Tuyere: enabling scalable memory workloads for system exploration
abstract
Memory technologies are under active development. Meanwhile, workloads on contemporary computing systems are increasing rapidly in size and diversity. Such dynamics in hardware and software further widen the gap between memory system design and performance evaluation. In this work, we propose a data-centric abstraction of high-performance computing applications for fast exploration of new memory technologies. We also provide a framework that uses a formal modeling language to describe the abstraction, automatically translates abstractions into memory traffic, and directly interfaces with cycle-accurate simulators. We evaluated the framework using 20 workloads and validated the memory traffic profile, the simulation results, and the relative memory changes of four memory technologies. Our results show that the data-centric abstraction can accurately capture application behavior adaptable to different input problems and can expedite system exploration.
Ivy Bo Peng, Jeffrey S. Vetter, Shirley V. Moore, Seyong Lee
HPDC1
2018 Siena: exploring the design space of heterogeneous memory systems
Ivy Bo Peng, Jeffrey S. Vetter
SC1
2018 Characterizing the performance benefit of hybrid memory system for HPC applications
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Jeffrey S. Vetter, Pietro Cicotti, Erwin Laure, Stefano Markidis
Parallel Comput.1
2018 MPI windows on storage for HPC applications
Sergio Rivas-Gomez, Roberto Gioiosa, Ivy Bo Peng, Gokcen Kestor, Sai Narasimhamurthy, Erwin Laure, Stefano Markidis
Parallel Comput.3
2017 Extending Message Passing Interface Windows to Storage
abstract
This paper presents an extension to MPI supporting the one-sided communication model and window allocations in storage. Our design transparently integrates with the current MPI implementations, enabling applications to target MPI windows in storage, memory or both simultaneously, without major modifications. Initial performance results demonstrate that the presented MPI window extension could potentially be helpful for a wide-range of use-cases and with low-overhead.
Sergio Rivas-Gomez, Stefano Markidis, Ivy Bo Peng, Erwin Laure, Gokcen Kestor, Roberto Gioiosa
CCGrid3
2017 Preparing HPC Applications for the Exascale Era: A Decoupling Strategy
abstract
Production-quality parallel applications are often a mixture of diverse operations, such as computation- and communication-intensive, regular and irregular, tightly coupled and loosely linked operations. In conventional construction of parallel applications, each process performs all the operations, which might result inefficient and seriously limit scalability, especially at large scale. We propose a decoupling strategy to improve the scalability of applications running on large-scale systems. Our strategy separates application operations onto groups of processes and enables a dataflow processing paradigm among the groups. This mechanism is effective in reducing the impact of load imbalance and increases the parallel efficiency by pipelining multiple operations. We provide a proof-of-concept implementation using MPI, the de-facto programming system on current supercomputers. We demonstrate the effectiveness of this strategy by decoupling the reduce, particle communication, halo exchange and I/O operations in a set of scientific and data-analytics applications. A performance evaluation on 8,192 processes of a Cray XC40 supercomputer shows that the proposed approach can achieve up to 4x performance improvement.
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Erwin Laure, Stefano Markidis
ICPP1
2017 RTHMS: a tool for data placement on hybrid memory system
abstract
Traditional scientific and emerging data analytics applications require fast, power-efficient, large, and persistent memories. Combining all these characteristics within a single memory technology is expensive and hence future supercomputers will feature different memory technologies side-by-side. However, it is a complex task to program hybrid-memory systems and to identify the best object-to-memory mapping. We envision that programmers will probably resort to use default configurations that only require minimal interventions on the application code or system settings. In this work, we argue that intelligent, fine-grained data placement can achieve higher performance than default setups.
Ivy Bo Peng, Roberto Gioiosa, Gokcen Kestor, Pietro Cicotti, Erwin Laure, Stefano Markidis
ISMM1
2017 Efficient alarm behavior analytics for telecom networks
Caifeng He, Guangjian Tian, Ivy Bo Peng, Jia Xing, Xiangbing Ruan, Haoran Xie 0001, Fu Lee Wang
Inf. Sci.5
2015 Evaluation of Parallel Communication Models in Nekbone, a Nek5000 Mini-Application
abstract
Nekbone is a proxy application of Nek5000, a scalable Computational Fluid Dynamics (CFD) code used for modelling incompressible flows. The Nekbone mini-application is used by several international co-design centers to explore new concepts in computer science and to evaluate their performance. We present the design and implementation of a new communication kernel in the Nekbone mini-application with the goal of studying the performance of different parallel communication models. First, a new MPI blocking communication kernel has been developed to solve Nekbone problems in a three-dimensional Cartesian mesh and process topology. The new MPI implementation delivers a 13% performance improvement compared to the original implementation. The new MPI communication kernel consists of approximately 500 lines of code against the original 7,000 lines of code, allowing experimentation with new approaches in Nekbone parallel communication. Second, the MPI blocking communication in the new kernel was changed to the MPI non-blocking communication. Third, we developed a new Partitioned Global Address Space (PGAS) communication kernel, based on the GPI-2 library. This approach reduces the synchronization among neighbor processes and is on average 3% faster than the new MPI-based, non-blocking, approach. In our tests on 8,192 processes, the GPI-2 communication kernel is 3% faster than the new MPI non-blocking communication kernel. In addition, we have used the OpenMP in all the versions of the new communication kernel. Finally, we highlight the future steps for using the new communication kernel in the parent application Nek5000.
Ilya Ivanov, Dana Akhmetova, Ivy Bo Peng, Stefano Markidis, Erwin Laure, Mirko Rahn, Valeria Bartsch, Alistair Hart, Paul F. Fischer
CLUSTER4
2015 The Cost of Synchronizing Imbalanced Processes in Message Passing Systems
abstract
Synchronization in message passing systems is achieved by communication among processes. System and architectural noise and different workloads cause processes to be imbalanced and to reach synchronization points at different time. Thus, both communication and imbalance impact the synchronization performance. In this paper, we study the algorithmic properties that allow the communication in synchronization to absorb the initial imbalance among processes. We quantify the imbalance absorption properties of different barrier algorithms using a LogP Monte Carlo simulator. We found that linear and f-way tournament barriers can absorb up to 95% of random exponential imbalance with the standard deviation equal to the communication time for one message. Dissemination, butterfly and pairwise exchange barriers, on the other hand, do not absorb imbalance but can effectively bound the post-barrier imbalance. We identify that synchronization transits from communication-dominated to imbalance-dominated when the standard deviation of imbalance distribution is more than twice the communication time for one message. In our study, f-way tournament barriers provided the best imbalance absorption rate and convenient communication time.
Ivy Bo Peng, Stefano Markidis, Erwin Laure
CLUSTER1