Stepan Nassyr

dblp:278/0602 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-0035-244XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Performance Portable BLAS3 Micro-Kernel Generator
Stepan Nassyr, Daniel Seibel, Prateek Chawla, Jayesh Badwaik, Andreas Herten, Dirk Pleiter
Euro-Par (1)1
2025 Performance-Portable Optimization and Analysis of Multiple Right-Hand Sides in a Lattice QCD Solver
abstract
Managing the high computational cost of iterative solvers for sparse linear systems is a known challenge in scientific computing. Moreover, scientific applications often face memory bandwidth constraints, making it critical to optimize data locality and enhance the efficiency of data transport. We extend the lattice QCD solver DD-$\alpha$AMG to incorporate multiple right-hand sides (rhs) for both the Wilson-Dirac operator evaluation and the GMRES solver, with and without odd-even preconditioning. To optimize auto-vectorization, we introduce a flexible interface that supports various data layouts and implement a new data layout for better SIMD utilization. We evaluate our optimizations on both x86 and Arm clusters, demonstrating performance portability with similar speedups. A key contribution of this work is the performance analysis of our optimizations, which reveals the complexity introduced by architectural constraints and compiler behavior. Additionally, we explore different implementations leveraging a new matrix instruction set for Arm called SME and provide an early assessment of its potential benefits.
Shiting Long, Gustavo Ramirez-Hidalgo, Stepan Nassyr, Jose Jimenez-Merchan, Andreas Frommer, Dirk Pleiter
HiPC3
2025 tfQMRgpu: a GPU-accelerated linear solver with block-sparse complex result matrix
abstract
Abstract We present , a GPU-accelerated iterative linear solver based on the transpose-free quasi-minimal residual (tfQMR) method. Designed for large-scale electronic structure calculations, particularly in the context of Korringa–Kohn–Rostoker density functional theory, efficiently handles block-sparse complex matrices arising from multiple scattering theory. The solver exploits GPU parallelism to accelerate convergence while leveraging memory-efficient sparse storage formats. By unifying the solution of multiple right-hand side (RHS) block vectors, significantly improves throughput, demonstrating up to a $$3.5\times$$ 3.5 × speedup on modern GPUs. Additionally, we introduce a flexible implementation framework that supports both explicit matrix-based and matrix-free operator formulations, such as high-order finite-difference stencils for real-space grid-based Green function calculations. Benchmarks on various NVIDIA GPUs demonstrate the solver’s efficiency, in some cases achieving over 56% of peak floating-point performance for block-sparse matrix multiplications. is open-source, providing interfaces for C, C++, Fortran, Julia, and Python, making it a versatile tool for high-performance computing applications that can benefit from the unification of RHS problems.
Paul F. Baumeister, Stepan Nassyr
J. Supercomput.2
2024 Exploring Processor Micro-architectures Optimised for BLAS3 Micro-kernels
Stepan Nassyr, Dirk Pleiter
Euro-Par (2)1
2024 Application-Driven Exascale: The JUPITER Benchmark Suite
abstract
Benchmarks are essential in the design of modern HPC installations, as they define key aspects of system components. Beyond synthetic workloads, it is crucial to include real applications that represent user requirements into benchmark suites, to guarantee high usability and widespread adoption of a new system. Given the significant investments in leadership-class supercomputers of the exascale era, this is even more important and necessitates alignment with a vision of Open Science and reproducibility. In this work, we present the JUPITER Benchmark Suite, which incorporates 16 applications from various domains. It was designed for and used in the procurement of JUPITER, the first European exascale supercomputer. We identify requirements and challenges and outline the project and software infrastructure setup. We provide descriptions and scalability studies of selected applications and a set of key takeaways. The JUPITER Benchmark Suite is released as open source software with this work at github.com/FZJ-JSC/jubench
Andreas Herten, Sebastian Achilles, Damian Alvarez, Jayesh Badwaik, Eric Behle, Mathis Bode, Thomas Breuer, Daniel Caviedes-Voullième, Mehdi Cherti, Adel Dabah, Salem El Sayed, Wolfgang Frings, Ana Gonzalez-Nicolas, Eric B. Gregory, Kaveh Haghighi Mood, Thorsten Hater, Jenia Jitsev, Chelsea Maria John, Jan H. Meinke, Catrin I. Meyer, Pavel Mezentsev, Jan-Oliver Mirus, Stepan Nassyr, Carolin Penke, Manoel Römmer, Ujjwal Sinha, Benedikt von St. Vieth, Olaf Stein, Estela Suarez, Dennis Willsch, Ilya Zhukov
SC23
2020 Porting Applications to Arm-based Processors
abstract
Arm-based server processors are becoming increasingly used for building massively parallel HPC systems. This triggers the need for porting HPC applications to such architectures and to collect more knowledge about performance benefits and challenges. In this contribution, we report on experiences made during our ongoing efforts to port applications to Arm-based platforms at the Jülich Supercomputing Centre (JSC). The performance of these applications is explored on different Arm-based node architectures plus an x86-based architecture for reference.
Bine Brank, Stepan Nassyr, Fatemeh Pouyan, Dirk Pleiter
CLUSTER2