Zane Fink

dblp:243/2560 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-1882-6578ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Embracing Dynamism: Control State Serialization for High-Performance Python
abstract
Python's ease of use has driven its adoption in data science, machine learning, and increasingly, high-performance computing (HPC), but its performance lags due to its dynamic nature. While many efforts accelerate Python by restricting its features, this work leverages Python's dynamism to accelerate it. We introduce control state serialization to Python through Sauerkraut, a library that captures the complete execution state (call stack, instruction pointers, operand stacks, local variables, global context) of running functions, complementing existing data serialization. Sauerkraut enables snapshots of function execution to be serialized, transferred, and resumed later or elsewhere. Sauerkraut is compatible with off-the-shelf Python installations. We demonstrate its utility by building Kombucha, a partial MPI implementation that provides general-purpose load balancing by migrating virtualized MPI ranks between processes using Sauerkraut to transfer control state. Evaluations show Sauerkraut adds minimal overhead over standard serialization, and Kombucha achieves significant speedups (up to 2.21x for a CPU Particle-in-Cell code and$1.43 \times$for a GPU Jacobi3D code with synthetic imbalance) with minimal application code changes. Control state serialization opens new avenues for performance optimization in Python, including load balancing, checkpoint/restart, and replay debugging, enhancing Python's suitability for the HPC community.
Zane Fink, Laxmikant V. Kalé
HiPC1
2024 HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
abstract
Recent advancements in Machine Learning (ML) have substantially improved its predictive and computational abilities, offering promising opportunities for surrogate modeling in scientific applications. By accurately approximating complex functions with low computational cost, ML-based surrogates can accelerate scientific applications by replacing computationally intensive components with faster model inference. However, integrating ML models into these applications remains a significant challenge, hindering the widespread adoption of ML surrogates as an approximation technique in modern scientific computing. We propose an easy-to-use directive-based programming model that enables developers to seamlessly describe the use of ML models in scientific applications. The runtime support, as instructed by the programming model, performs data assimilation using the original algorithm and can replace the algorithm with model inference. Our evaluation across five benchmarks, testing over 5000 ML models, shows up to $83.6 \times$ speed improvements with minimal accuracy loss (as low as 0.01 RMSE).
Zane Fink, Konstantinos Parasyris, Praneet Rathi, Giorgis Georgakoudis, Harshitha Menon, Peer-Timo Bremer
SC1
2023 HPAC-Offload: Accelerating HPC Applications with Portable Approximate Computing on the GPU
abstract
The end of Dennard scaling and the slowdown of Moore's law led to a shift in technology trends towards parallel architectures, particularly in HPC systems. To continue providing performance benefits, HPC should embrace Approximate Computing (AC), which trades application quality loss for improved performance. However, existing AC techniques have not been extensively applied and evaluated in state-of-the-art hardware architectures such as GPUs, the primary execution vehicle for HPC applications today.
Zane Fink, Konstantinos Parasyris, Giorgis Georgakoudis, Harshitha Menon
SC1
2022 Accelerating communication for parallel programming models on GPU systems
Zane Fink, Sam White, Nitin Bhat, David F. Richards, Laxmikant V. Kalé
Parallel Comput.2
2019 Accelerating the Unacceleratable: Hybrid CPU/GPU Algorithms for Memory-Bound Database Primitives
abstract
Many database operations have a low compute to memory access ratio. In heterogeneous systems, where a graphics processing unit (GPU) is interconnected via PCIe, the data transfer bottleneck is perceived as insurmountable to achieving performance gains on these memory-bound database primitives. On the other hand, several compute-bound database operations have been shown to achieve significant performance gains using the GPU. This leads to CPU-only memory-bound applications having an increasingly non-negligible impact on database query throughput. In this paper we examine several of these overlooked algorithms, including (i) batched predecessor searches; (ii) multiway merging; and, (iii) partitioning. We examine the performance of parallel CPU-only, GPU-only, and hybrid CPU/GPU approaches, and show that hybrid algorithms achieve respectable performance gains. We develop a model that considers main memory accesses and PCIe data transfers, which are two major bottlenecks for hybrid CPU/GPU algorithms. The model lets us analytically determine how to distribute work between the CPU and GPU to maximize resource utilization while minimizing load imbalance. We show that our model can accurately predict the fraction of work to be sent to each architecture, and consequently, confirms that these overlooked database primitives can be accelerated despite their memory-bound nature.
Michael G. Gowanlock, Benjamin Karsin, Zane Fink, Jordan Wright
DaMoN3