Katarzyna Swirydowicz

dblp:157/8492 · also Kasia Swirydowicz · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-5758-5394ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization
Katarzyna Swirydowicz, Nicholson Koukpaizan, Maksudul Alam, Shaked Regev, Michael A. Saunders, Slaven Peles
Parallel Comput.1
2024 FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
abstract
NVIDIA Tensor Cores and AMD Matrix Cores (together called Matrix Accelerators) are of growing interest in high-performance computing and machine learning owing to their high performance. Unfortunately, some of their crucial numerical attributes pertaining to departures from full IEEE floating-point compatibility are not documented. This makes it impossible to reliably port codes across these differing accelerators. This paper contributes a collection of Feature Targeted Tests for Numerical Properties that that help determine these features across five floating-point formats, four rounding modes and additional that highlight the rounding behaviors and preservation of extra precision bits. To show the practical relevance of FTTN, we design a simple matrix-multiplication test designed with insights gathered from our feature-tests. We executed this very simple test on five platforms, producing different answers: V100, A100, and MI250X produced 0, MI100 produced 255.875, and Hopper H100 produced 191.875. Our matrix multiplication tests employ patterns found in iterative refinement-based algorithms, highlighting the need to check for significant result variability when porting code across GPUs.
Ang Li 0006, Bo Fang 0002, Katarzyna Swirydowicz, Ignacio Laguna, Ganesh Gopalakrishnan
CCGrid4
2024 Discovery of Floating-Point Differences Between NVIDIA and AMD GPUs
abstract
NVIDIA and AMD GPUs are fundamental components in contemporary high-performance systems, boosting computational capabilities in the HPC and AI fields.However, a clear understanding of the nuances in floating-point operations between these GPU variants is crucial to avoid introducing errors during software development or porting, and such clarity is currently insufficient.The complexity of this issue is amplified when considering the variety of floating-point precision options (such as FP16, FP32, etc.), floating-point formats (like standard floats, bfloats, etc.), and the different execution units (elementary units, matrix/tensor cores, etc.).As it stands, much of this information is either not well-known or is difficult to obtain. Our work aims to shed light on these areas through a pioneering testing-guided methodology that seeks to unravel many of these uncertainties.We are in the process of developing a series of tests that uncover the numerical discrepancies in elementary computing units, the built-in math libraries, and the numerical properties of matrix accelerators present in both NVIDIA (tensor cores) and AMD GPUs (matrix cores).The significance of this testing approach extends beyond current GPU models; it is designed to be forward-compatible with upcoming GPU technologies. We have already identified discrepancies as significant as 7 ulps for trigonometric functions at FP32 precision and 3 ulps at FP64 precision between NVIDIA and AMD GPUs. Additionally, our comprehensive examination has documented the behaviors of matrix cores (NVIDIA) and tensor cores (AMD), including their rounding modes (such as truncation and round-to-nearest), the extent of extra internal bits maintained (specifically, whether an additional 3 bits are retained), the handling of subnormal numbers in inputs and outputs and the FMA features in these units. This analysis spans four distinct floating-point formats and multiple GPU models, including NVIDIA’s V100, A100, H100 and AMD’s MI100 and MI250X.We believe that the information now being disclosed will reduce the risk of porting errors when codes are adapted across these different hardware platforms.
Ang Li 0006, Bo Fang 0002, Katarzyna Swirydowicz, Ignacio Laguna, Ganesh Gopalakrishnan
CCGrid4
2024 A Performance and Energy Study of GPU-Resident Preconditioners for Conjugate Gradient Solvers: In the Context of Existing and Novel Approaches
abstract
Optimizing a particular subprogram out of the set of Basic (sparse) Linear Algebra Subprograms (BLAS) for a given architecture is a common topic of research. In applications, however, these BLAS functions rarely appear in isolation; usually, many of them are used together, in various combinations and with varying inputs. As the need to solve a large, sparse linear system is ubiquitous throughout HPC applications, linear solvers constitute a realistic, sufficiently complex and well-defined representative use case for composite BLAS routines. To this end, based on a representative set of matrices drawn from a diverse set of fields, we present a framework to study, from the performance and energy perspective, the efficacy of GPU-resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process. The development of this preconditioner was motivated by solving large graph Laplacian linear systems, for which the existing preconditioners either perform slow on GPU-based platforms or are not applicable. We compare the performance of these preconditioners on different hardware accelerator architectures, i.e., AMD MI250X, MI100, Nvidia A100, V100, and Jetson. Our experiments reveal performance trade-offs and provide information on how to select the best strategy for the given linear system, dictated by its properties, and the platform of interest. We demonstrate the application of our novel preconditioner for solving CG and graph Laplacian systems. Overall, the framework can be utilized as a benchmark to guide informed decisions in choosing a specific preconditioner, i.e., whether it is better to rely on the performance of a triangular solver or on the performance of sparse matrix-vector product. Finally, by considering power consumption to solve the linear systems, we report the energy footprint for the solvers.
Katarzyna Swirydowicz, Jesun Sahariar Firoz, Joseph B. Manzano, Mahantesh Halappanavar, Kevin J. Barker
SBAC-PAD1
2023 Design and Evaluation of GPU-FPX: A Low-Overhead tool for Floating-Point Exception Detection in NVIDIA GPUs
abstract
Floating-point exceptions occurring during numerical computations can be a serious threat to the validity of the computed results if they are not caught and diagnosed Unfortunately, on NVIDIA GPUs-today's most widely used types and which do not have hardware exception traps-this task must be carried out in software. Given the prevalence of closed-source kernels, efficient binary-level exception tracking is essential. It is also important to know how exceptions flow through the code, whether they alter the code behavior and additionally whether these exceptions can be detected at the program outputs or are killed inside program flow-paths.
Ignacio Laguna, Bo Fang 0002, Katarzyna Swirydowicz, Ang Li 0006, Ganesh Gopalakrishnan
HPDC4
2022 Low-synch Gram-Schmidt with delayed reorthogonalization for Krylov solvers
Daniel Bielich, Julien Langou, Stephen J. Thomas, Katarzyna Swirydowicz, Ichitaro Yamazaki, Erik G. Boman
Parallel Comput.4
2022 Linear solvers for power grid optimization problems: A review of GPU-accelerated linear solvers
Katarzyna Swirydowicz, Eric Darve, Wesley B. Jones, Jonathan Maack, Shaked Regev, Michael A. Saunders, Stephen J. Thomas, Slaven Peles
Parallel Comput.1