VLDB 2026 Research / reviewers in the wild / expert
Hal Finkel
dblp:121/2175
· DBLP profile ↗
23ranked-venue papers
0as first author
3since 2021 · last 2022
0000-0002-7551-7122ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 4Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimizationabstractAbstract We develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat‐3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd–Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd–Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application. Xingfu Wu, Michael Kruse, Prasanna Balaprakash, Hal Finkel, Paul D. Hovland, Valerie Taylor 0001, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | OpenMP application experiences: Porting to accelerated nodes
Seonmyeong Bak, Colleen Bertoni, Swen Böhm, Reuben D. Budiardja, Barbara M. Chapman, Johannes Doerfert, Markus Eisenbach 0002, Hal Finkel, Oscar R. Hernandez, Joseph Huber, Shintaro Iwasaki, Vivek Kale, Paul R. C. Kent, JaeHyuk Kwack, Meifeng Lin, Piotr Luszczek, Ye Luo 0001, Buu Pham, Swaroop Pophale, Kiran Ravikumar, Vivek Sarkar, Thomas Scogland, Shilei Tian, P. K. Yeung |
Parallel Comput. | 8 |
| 2021 | Extending C++ for Heterogeneous Quantum-Classical ComputingabstractWe present qcor—a language extension to C++ and compiler implementation that enables heterogeneous quantum-classical programming, compilation, and execution in a single-source context. Our work provides a first-of-its-kind C++ compiler enabling high-level quantum kernel (function) expression in a quantum-language agnostic manner, as well as a hardware-agnostic, retargetable compiler workflow targeting a number of physical and virtual quantum computing backends. qcor leverages novel Clang plugin interfaces and builds upon the XACC system-level quantum programming framework to provide a state-of-the-art integration mechanism for quantum-classical compilation that leverages the best from the community at-large. qcor translates quantum kernels ultimately to the XACC intermediate representation, and provides user-extensible hooks for quantum compilation routines like circuit optimization, analysis, and placement. This work details the overall architecture and compiler workflow for qcor, and provides a number of illuminating programming examples demonstrating its utility for near-term variational tasks, quantum algorithm expression, and feed-forward error correction schemes. Alex McCaskey, Thien Nguyen 0001, Anthony Santana, Daniel Claudino, Tyler Kharazi, Hal Finkel |
ACM Trans. Quantum Comput. | 6 |
| 2019 | Base64 Encoding on Heterogeneous Computing PlatformsabstractBase64 encoding has many applications on the Web. Previous studies investigated the optimizations of Base64 encoding algorithm on central processing units (CPUs). In this paper, we describe the optimizations of the algorithm on heterogeneous computing platforms. More specifically, we explain the algorithm, convert the algorithm to kernels written in CUDA C/C++ and Open Computing Language (OpenCL), optimize the CUDA and OpenCL applications with CUDA and OpenCL streams which can overlap data transfers with kernel computations, and vectorize the CUDA and OpenCL kernels to improve kernel throughput. We evaluate the impact of the number of streams upon the kernel performance on an NVIDIA Pascal P100 graphics processing unit (GPU) and a Nallatech 385A card that features an Intel Arria 10 GX1150 field-programmable gate array (FPGA). We also measure the performance and power of the applications on the CPU, GPU, and FPGA to know the advantage of each platform and the benefit of kernel offloading. The experiments show that using vector data types in the kernels is not for performance, and more work-items is better than large vectors per work-item on the GPU. OpenCL and CUDA streams can achieve almost the same performance on the GPU, but streams should be used with caution when GPU resources are underutilized. On the FPGA, kernel vectorization using 16 vector lanes can achieve the highest performance when the number of streams is one. However, increasing the vector width per work-item and the number of streams can decrease the kernel computation time for each stream, and thereby reduce the number of concurrent operations across the streams. While the raw performance on the GPU is 3.1X higher than that on the FPGA, the FPGA consumes 3.4X less power. A comparison with a state-of-the-art implementation on an Intel CPU server shows an increasing benefit of kernel offloading. Zheming Jin, Hal Finkel |
ASAP | 2 |
| 2019 | Evaluation of Medical Imaging Applications using SYCLabstractAs opposed to the Open Computing Language (OpenCL) programming model in which host and device codes are written in different languages, the SYCL programming model can combine host and device codes for an application in a type-safe way to improve development productivity. In this paper, we chose two medical imaging applications (Heart Wall and Particle Filter) in the Rodinia benchmark suite to study the performance and programming productivity of the SYCL programming model. More specifically, we introduced the SYCL programming model, shared our experience of implementing the applications using SYCL, and compared the performance and programming portability of the SYCL implementations with the OpenCL implementations on an Intel® Xeon® CPU and an Iris® Pro integrated GPU. The results are promising. For the Heart Wall application, the SYCL implementation is on average 15% faster than the OpenCL implementation on the GPU. For the Particle Filter application, the SYCL implementation is 3% slower than the OpenCL implementation on the GPU, but it is 75% faster on the CPU. Using lines of code as an indicator of programming productivity, the SYCL host program reduces the lines of code of the OpenCL host program by 52% and 38% for the Heart Wall and Particle Filter applications, respectively. Zheming Jin, Hal Finkel |
BIBM | 2 |
| 2019 | Exploration of OpenCL 2D Convolution Kernels on Intel FPGA, CPU, and GPU PlatformsabstractThere is a need to evaluate the resource usage and optimize the performance of the floating-point 2D convolution kernels on a recent FPGA which features large numbers of hardened floating-point digital signal processing blocks and an increasingly large on-chip memory. In this paper, we presented an OpenCL 2D convolution kernel with configurable parameters for specifying the precision, sizes of filter and block, vectorization width, and compute-unit duplication factor. Then, we instantiated a set of specific instances of the kernel with a fixed filter size to evaluate their resource usage and performance to narrow down the exploration space. Based on the evaluation results on an Intelo Arria 10 FPGA using high-level synthesis, we evaluated the kernels with different filter sizes within the pruned exploration space. Compared to the baseline implementation in which the vectorization width is two and the block size is 32$\times$ 32, our optimizations improve the performance by a factor ranging from 1. 9X to 3X for the single-precision kernels, and from 2. 2X to 3. 37X for the half-precision kernels. Furthermore, we evaluated the performance and power of the kernels on an Intel®Xeon®CPU and an IrisTMPro integrated GPU. We found that the FPGA could achieve the highest performance for a 9$\times$ 9 filter among the CPU, GPU, and FPGA, but the GPU can achieve the highest performance for other filter sizes. Zheming Jin, Hal Finkel |
IEEE BigData | 2 |
| 2019 | A Case Study of k-means Clustering using SYCLabstractAs opposed to the OpenCL programming model in which host and device codes are written in two programming languages, the SYCL programming model combines them for an application in a type-safe way to improve development productivity. As a popular cluster analysis algorithm, k-means has been implemented using programming models such as OpenMP, OpenCL, and CUDA. Developing a SYCL implementation of k-means as a case study allows us to have a better understanding of performance portability and programming productivity of the SYCL programming model. Specifically, we explained the k-means benchmark in Rodinia, described our efforts of porting the OpenCL k-means benchmark, and evaluated the performance of the OpenCL and SYCL implementations on the Intel®Haswell, Broadwell, and Skylake processors. We summarized the migration steps from OpenCL to SYCL, compiled the SYCL program using Codeplay and Intel®SYCL compilers, analyzed the SYCL and OpenCL programs using an open-source profiling tool which can intercept OpenCL runtime calls, and compared the performance of the implementations on Intel®CPUs and integrated GPU. The experimental results show that the SYCL version in which the kernels run on the GPU is 2% and 8% faster than the OpenCL version for the two large datasets. However, the OpenCL version is still much faster than the SYCL version on the CPUs. Compared to the Intel®Haswell and Skylake CPUs, running the k-means benchmark on the Intel®Broadwell low-power processor with a CPU and an integrated GPU can achieve the lowest energy consumption. In terms of programming productivity, the lines of code of the SYCL program are 51% fewer than those of the OpenCL program. Zheming Jin, Hal Finkel |
IEEE BigData | 2 |
| 2019 | Accelerating Hyperdimensional Classifier on Multiple GPUsabstractAmong brain-inspired computing paradigms, hyperdimensional (HD) computing is based on mathematical properties of high-dimensional spaces which show remarkable agreement with brain-controlled behaviors [1] . In [2] , the authors present an HD classifier for the task of identifying the language of text samples based on letter N -grams. They describe a computing architecture in which an encoding module generates a hypervector for each text sample and a search module compares the generated vector with a set of trained hypervectors. They provide an open-source implementation of their HD classifier written in hardware description language (HDL). In addition, they implemented the classifier with the 65-nm technology library, and evaluated the efficiency and accuracy of the classifier. Zheming Jin, Hal Finkel |
CLUSTER | 2 |
| 2019 | Exploring the Random Network of Hodgkin and Huxley Neurons with Exponential Synaptic Conductances on OpenCL FPGA PlatformabstractWe choose a random network of Hodgkin-Huxley (HH) neurons with exponential synaptic conductance as a study of accelerating the simulation of networks of spiking neurons on an FPGA. Focused on the conductance-based HH (COBAHH) benchmark, we execute the benchmark on a general-purpose simulator for spiking neural networks, identify a computationally intensive kernel in the generated C++ code, convert the kernel to a portable OpenCL kernel, and describe the kernel optimizations which can reduce the resource utilizations and improve the kernel performance. We evaluate the kernel on an Intel Arria 10 based FPGA platform, an Intel Xeon 16-core CPU, and an NVIDIA Tesla P100 GPU. FPGAs are promising for the simulation of spiking neuron network. Zheming Jin, Hal Finkel |
FCCM | 2 |
| 2019 | OpenCL Kernel Vectorization on the CPU, GPU, and FPGA: A Case Study with Frequent Pattern CompressionabstractOpenCL promotes code portability, and natively supports vectorized data types, which allows developers to potentially take advantage of the single-instruction-multiple-data instructions on CPUs, GPUs, and FPGAs. FPGAs are becoming a promising heterogeneous computing component. In our study, we choose a kernel used in frequent pattern compression as a case study of OpenCL kernel vectorizations on the three computing platforms. We describe different pattern matching approaches for the kernel, and manually vectorize the OpenCL kernel by a factor ranging from 2 to 16. We evaluate the kernel on an Intel Xeon 16-core CPU, an NVIDIA P100 GPU, and a Nallatech 385A FPGA card featuring an Intel Arria 10 GX1150 FPGA. Compared to the optimized kernel that is not vectorized, our vectorization can improve the kernel performance by a factor of 16 on the FPGA. The performance improvement ranges from 1 to 11.4 on the CPU, and from 1.02 to 9.3 on the GPU. The effectiveness of kernel vectorization depends on the work-group size. Zheming Jin, Hal Finkel |
FCCM | 2 |
| 2019 | Base64 Encoding on OpenCL FPGA PlatformabstractBase64 encoding has many applications on the Web. Previous studies are focused on improving the efficiency of Base64 encoding on central processing units (CPUs). As field-programmable gate arrays (FPGAs) are becoming promising heterogeneous computing components in high-performance computing (HPC), and high-level synthesis (HLS) is more mature, we are motivated to optimize Base64 encoding on an FPGA using HLS. In this paper, we explain the algorithm, converts the algorithm to a kernel written in Open Computing Language (OpenCL), and optimize the kernel targeting an Intel Arria 10 FPGA. We evaluate the performance and power of the kernel implementations on the CPU, graphics processing units (GPUs), and FPGA computing platforms. The experimental results show that we can significantly improve the performance of Base64 encoding with the FPGA-specific optimizations. Compared to an Intel Xeon Platinum 8167 CPU, an Nvidia Tesla K80 GPU, and an Nvidia Tesla P100 GPU, the performance (the number of cycles per byte) of Base64 encoding on an Arria10-based FPGA platform is 3.98X higher than that on the K80 GPU, 17X higher than that on the CPU, and 1.83X lower than that on the P100 GPU for large input data sizes. The performance per watt on the FPGA is 1.1X lower than that on the P100 GPU, and 8.25X and 13.2X higher than that on the CPU and the K80 GPU, respectively. Zheming Jin, Hal Finkel |
FPGA | 2 |
| 2019 | Nuclear Reactor Simulations on OpenCL FPGA PlatformabstractField-programmable gate arrays (FPGAs) are becoming a promising choice as a heterogeneous computing component for scientific computing when floating-point optimized architectures are added to the current FPGAs. The maturing high-level synthesis (HLS) tools, such as Intel FPGA SDK for OpenCL, provide a streamlined design flow to facilitate parallel application on FPGAs. In this paper, we evaluate and optimize the OpenCL implementations of three nuclear reactor simulation applications (XSBench, RSBench, and SimpleMOC kernel) on a heterogeneous computing platform that consists of a general-purpose CPU and an FPGA. We introduce the applications, and describe their OpenCL implementations and optimization methods on an Arria10-based FPGA platform. Compared with the baseline kernel implementations, our optimizations increase the performance of the three kernels by a factor of 35, 295, and 102, respectively. We compare the performance, power, and performance per watt of the three applications on an Intel Xeon 16-core CPU, an Nvidia Tesla K80 GPU, and an Intel Arria10 GX1150 FPGA. The performance per watt on the FPGA is competitive. For XSBench, the performance per watt on the FPGA is 1.43X higher than that on the CPU, and 2.58X lower than that on the GPU. For RSBench, the performance per watt on the FPGA is 3.6X higher than that on the CPU, and 5.8X lower than that on the GPU. For SimpleMOC kernel, the performance per watt on the FPGA is 1.74X higher than that on the CPU, and 1.65X lower than that on the GPU. Zheming Jin, Hal Finkel |
FPGA | 2 |
| 2019 | Full-state quantum circuit simulation by using data compressionabstractQuantum circuit simulations are critical for evaluating quantum algorithms and machines. However, the number of state amplitudes required for full simulation increases exponentially with the number of qubits. In this study, we leverage data compression to reduce memory requirements, trading computation time and fidelity for memory space. Specifically, we develop a hybrid solution by combining the lossless compression and our tailored lossy compression method with adaptive error bounds at each timestep of the simulation. Our approach optimizes for compression speed and makes sure that errors due to lossy compression are uncorrelated, an important property for comparing simulation output with physical machines. Experiments show that our approach reduces the memory requirement of simulating the 61-qubit Grover's search algorithm from 32 exabytes to 768 terabytes of memory on Argonne's Theta supercomputer using 4,096 nodes. The results suggest that our techniques can increase the simulation size by 2~16 qubits for general quantum circuits. Xin-Chuan Wu, Sheng Di, Emma Maitreyee Dasgupta, Franck Cappello, Hal Finkel, Yuri Alexeev, Fred Chong |
SC | 5 |
| 2018 | Optimizing Radial Basis Function Kernel on OpenCL FPGA PlatformabstractIn this paper, we optimize a widely used kernel, radial basis function, in a support vector machine as a case study to evaluate the potential of using FPGAs and the capabilities of high-level synthesis (HLS) for data intensive applications. We explain the HLS flow, and use it to develop and evaluate the kernels optimized with vectorization, loop unrolling, and half-precision storage format. Our optimizations improve the kernel performance by a factor of 15.8 compared to a baseline kernel on the Nallatech 385A FPGA card that features an Intel Arria 10 GX 1150 FPGA. The half storage format can reduce the DSP and memory utilizations at the cost of increasing the logic utilization. Compared to the single-precision floating-point kernels, the half-precision kernels can reduce the dynamic power consumption on the FPGA by approximately 30%. In terms of energy efficiency, the performance per watt on the FPGA platform is approximately 3X higher than that on an Intel Xeon 16-core CPU, and 1.8X higher than that on an Nvidia Tesla K80 GPU. On the other hand, the raw performance on the FPGA is approximately 2X and 2.7X lower than that on the CPU and GPU, respectively. Zheming Jin, Hal Finkel |
IEEE BigData | 2 |
| 2018 | Bob Jenkins Lookup3 Hash Function on OpenCL FPGA PlatformabstractField-programmable gate array (FPGA) is a promising choice as a heterogeneous computing component for energy-aware and high-performance applications. Emerging high-level synthesis (HLS) tools such as Intel FPGA Software Development Kit for Open Computing Language (OpenCL) offer a streamlined design flow to facilitate the use of FPGAs for scientists and researchers. In this paper, we focus on the optimizations of the OpenCL design of the Bob Jenkins lookup3 hash function which is used in the open source software version of Memcached. We describe in details the optimizations of the kernel on the FPGA, and evaluate the resource utilizations, performance, and performance per watt of the kernel implementations on an Arria10-based FPGA platform. The experimental results show that the optimized design can achieve 3.46X speedup in kernel execution time compared to the baseline implementation on the Nallatech 385A FPGA card that features an Arria 10 GX 1150 FPGA chip. For the performance per watt, we achieve 8 MHash/watt on the Arria 10 FPGA, which is 14X and 1.2X improvement over an Intel Xeon E5 CPU and an Nvidia K80 GPU, respectively. Zheming Jin, Hal Finkel |
IEEE BigData | 2 |
| 2018 | Evaluation of OpenCL Performance-oriented Optimizations for Streaming Kernels on the FPGA: (Abstract Only)abstractThe streaming applications efficiently and High-level synthesis (HLS) tools allow people without complex hardware design knowledge to evaluate an application on FPGAs, there is an opportunity and a need to understand where OpenCL and FPGA can play in the streaming domains. To this end, we evaluate the overhead of the OpenCL infrastructure on the Nallatech 385A FPGA board that features an Arria 10 GX1150 FPGA. Then we explore the implementation space and discuss the performance optimization techniques for the streaming kernels using the OpenCL-to-FPGA HLS tool. On the target platform, the infrastructure overhead requires 12% of the FPGA memory and logic resources. The latency of the single work-item kernel execution is 11 us and the maximum frequency of a kernel implementation is around 300 MHz. The experimental results of the streaming kernels show FPGA resources, such as block RAMs and DSPs, can limit the kernel performance before the constraint of memory bandwidth takes effect. Kernel vectorization and compute unit duplication are practical optimization techniques that can improve the kernel performance by a factor of 2 to 10. The combination of the two techniques can achieve the best performance. To improve the performance of compute unit duplication, the local work size needs to be tuned and the optimal value can increase the performance by a factor of 3 to 70 compared to the default value. Zheming Jin, Hal Finkel |
FPGA | 2 |
| 2018 | Evaluating and Optimizing OpenCL Base64 Data Unpacking Kernel with FPGAabstractDevelopment of applications using OpenCL targeting FPGAs is an emerging approach on heterogeneous computing systems. This paper uses the data unpacking algorithm in Base64 encoding as a case study to present programming and optimization techniques, and experimental results of the OpenCL-based implementations on an FPGA. We explain the algorithm and evaluate the performance of the kernel implementations with Intel's FPGA OpenCL SDK. The experimental results show kernel vectorization and duplication are two optimization techniques that can improve the kernel performance. The performance of kernel duplication is also closely related to the local work size. Our experiment shows 16-lane vectorization increases the bandwidth by a factor of 2 to 10 for large input data sizes. Moreover, the performance of kernel duplication using 16 compute units is 40% to 1.5% less than that of kernel vectorization depending on the input size. Tuning the local work size can improve the kernel performance by a factor of 3 to 23. For this kernel, using local memory is not an effective technique to improve the kernel performance because input data is not reused. A combination of vectorization and duplication achieves the highest performance of 12.3 GiB/s. Compared to an Intel Xeon E5 CPU and an Nvidia Tesla K80 GPU, the performance of the kernel on the Arria 10 FPGA is 6.7X faster than the CPU and 3X slower than the GPU. The performance per watt on the FPGA is 20.5X higher than the CPU and 1.19X lower than the GPU. Zheming Jin, Iris Johnson, Hal Finkel |
PDP | 3 |
| 2017 | Evaluating irregular memory access on OpenCL FPGA platforms: A case study with XSBenchabstractFPGAs are becoming an attractive choice as a heterogeneous computing unit for scientific computing because FPGA vendors are adding floating-point-optimized architectures to their product lines. Additionally, high-level synthesis (HLS) tools such as Altera OpenCL SDK are emerging, which could potentially break the FPGA programming wall and provide a streamlined flow for domain experts in scientific computing. On the other hand, providing high performance in the presence of irregular memory access patterns to off-chip memory remains a challenge for the automated synthesis flows. In this paper, we study the performance/energy characteristics of OpenCL-generated FPGA designs on irregular memory access patterns, targeting XSBench, a memory-intensive Monte Carlo simulation code, as a case study. To complete our study, we implement XSBench in OpenCL and study optimization strategies for FPGAs. We observe that our OpenCL implantation of XSBench achieves 50 % higher energy efficiency on an Intel Arria10-based FPGA platform than that on an Intel Xeon 8-core CPU while trading off 35 % of performance. Yingyi Luo, Xianshan Wen, Kazutomo Yoshii, Seda Ogrenci Memik, Gokhan Memik, Hal Finkel, Franck Cappello |
FPL | 6 |
| 2017 | Trends in Data Locality Abstractions for HPC SystemsabstractThe cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems. Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2015 | Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approachabstractLarge-scale simulations can produce hundreds of terabytes to petabytes of data, complicating and limiting the efficiency of workflows. Traditionally, outputs are stored on the file system and analyzed in post-processing. With the rapidly increasing size and complexity of simulations, this approach faces an uncertain future. Trending techniques consist of performing the analysis in-situ, utilizing the same resources as the simulation, and/or off-loading subsets of the data to a compute-intensive analysis system. We introduce an analysis framework developed for HACC, a cosmological N-body code, that uses both in-situ and co-scheduling approaches for handling petabyte-scale outputs. We compare different analysis set-ups ranging from purely off-line, to purely in-situ to in-situ/co-scheduling. The analysis routines are implemented using the PISTON/VTK-m framework, allowing a single implementation of an algorithm that simultaneously targets a variety of GPU, multi-core, and many-core architectures. Christopher M. Sewell, Katrin Heitmann, Hal Finkel, George Zagaris, Suzanne Parete-Koon, Patricia K. Fasel, Adrian Pope, Nicholas Frontiere, Li-Ta Lo, O. E. Bronson Messer, Salman Habib 0002, James P. Ahrens |
SC | 3 |
| 2014 | Scalable Parallel I/O on a Blue Gene/Q Supercomputer Using Compression, Topology-Aware Data Aggregation, and SubfilingabstractIn this paper, we propose an approach to improving the I/O performance of an IBM Blue Gene/Q supercomputing system using a novel framework that can be integrated into high performance applications. We take advantage of the system's tremendous computing resources and high interconnection bandwidth among compute nodes to efficiently exploit I/O bandwidth. This approach focuses on lossless data compression, topology-aware data movement, and subfiling. The efficacy of this solution is demonstrated using microbenchmarks and an application-level benchmark. Huy Bui, Hal Finkel, Venkatram Vishwanath, Salman Habib 0002, Katrin Heitmann, Jason Leigh, Michael E. Papka, Kevin Harms |
PDP | 2 |
| 2013 | HACC: extreme scaling and performance across diverse architecturesabstractSupercomputing is evolving towards hybrid and accelerator-based architectures with millions of cores. The HACC (Hardware/Hybrid Accelerated Cosmology Code) framework exploits this diverse landscape at the largest scales of problem size, obtaining high scalability and sustained performance. Developed to satisfy the science requirements of cosmological surveys, HACC melds particle and grid methods using a novel algorithmic structure that flexibly maps across architectures, including CPU/GPU, multi/many-core, and Blue Gene systems. We demonstrate the success of HACC on two very different machines, the CPU/GPU system Titan and the BG/Q systems Sequoia and Mira, attaining unprecedented levels of scalable performance. We demonstrate strong and weak scaling on Titan, obtaining up to 99.2% parallel efficiency, evolving 1.1 trillion particles. On Sequoia, we reach 13.94 PFlops (69.2% of peak) and 90% parallel efficiency on 1,572,864 cores, with 3.6 trillion particles, the largest cosmological benchmark yet performed. HACC design concepts are applicable to several other supercomputer applications. Salman Habib 0002, Vitali A. Morozov, Nicholas Frontiere, Hal Finkel, Adrian Pope, Katrin Heitmann |
SC | 4 |
| 2012 | The universe at extreme scale: multi-petaflop sky simulation on the BG/QabstractRemarkable observational advances have established a compelling cross-validated model of the Universe. Yet, two key pillars of this model -- dark matter and dark energy -- remain mysterious. Next-generation sky surveys will map billions of galaxies to explore the physics of the 'Dark Universe'. Science requirements for these surveys demand simulations at extreme scales; these will be delivered by the HACC (Hybrid/Hardware Accelerated Cosmology Code) framework. HACC's novel algorithmic structure allows tuning across diverse architectures, including accelerated and multi-core systems. On the IBM BG/Q, HACC attains unprecedented scalable performance - currently 6.23 PFlops at 62% of peak and 92% parallel efficiency on 786,432 cores (48 racks) - at extreme problem sizes with up to almost two trillion particles, larger than any cosmological simulation yet performed. HACC simulations at these scales will for the first time enable tracking individual galaxies over the entire volume of a cosmological survey. Salman Habib 0002, Vitali A. Morozov, Hal Finkel, Adrian Pope, Katrin Heitmann, Kalyan Kumaran, Tom Peterka, Joseph A. Insley, David Daniel, Patricia K. Fasel, Nicholas Frontiere, Zarija Lukic |
SC | 3 |