EDBT 2026 Demo / reviewers in the wild / expert
Stepan Nassyr
dblp:278/0602
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-0035-244XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Portable BLAS3 Micro-Kernel Generator
Stepan Nassyr, Daniel Seibel, Prateek Chawla, Jayesh Badwaik, Andreas Herten, Dirk Pleiter |
Euro-Par (1) | 1 |
| 2025 | Performance-Portable Optimization and Analysis of Multiple Right-Hand Sides in a Lattice QCD SolverabstractManaging the high computational cost of iterative solvers for sparse linear systems is a known challenge in scientific computing. Moreover, scientific applications often face memory bandwidth constraints, making it critical to optimize data locality and enhance the efficiency of data transport. We extend the lattice QCD solver DD-$\alpha$AMG to incorporate multiple right-hand sides (rhs) for both the Wilson-Dirac operator evaluation and the GMRES solver, with and without odd-even preconditioning. To optimize auto-vectorization, we introduce a flexible interface that supports various data layouts and implement a new data layout for better SIMD utilization. We evaluate our optimizations on both x86 and Arm clusters, demonstrating performance portability with similar speedups. A key contribution of this work is the performance analysis of our optimizations, which reveals the complexity introduced by architectural constraints and compiler behavior. Additionally, we explore different implementations leveraging a new matrix instruction set for Arm called SME and provide an early assessment of its potential benefits. Shiting Long, Gustavo Ramirez-Hidalgo, Stepan Nassyr, Jose Jimenez-Merchan, Andreas Frommer, Dirk Pleiter |
HiPC | 3 |
| 2025 | tfQMRgpu: a GPU-accelerated linear solver with block-sparse complex result matrixabstractAbstract We present , a GPU-accelerated iterative linear solver based on the transpose-free quasi-minimal residual (tfQMR) method. Designed for large-scale electronic structure calculations, particularly in the context of Korringa–Kohn–Rostoker density functional theory, efficiently handles block-sparse complex matrices arising from multiple scattering theory. The solver exploits GPU parallelism to accelerate convergence while leveraging memory-efficient sparse storage formats. By unifying the solution of multiple right-hand side (RHS) block vectors, significantly improves throughput, demonstrating up to a $$3.5\times$$ 3.5 × speedup on modern GPUs. Additionally, we introduce a flexible implementation framework that supports both explicit matrix-based and matrix-free operator formulations, such as high-order finite-difference stencils for real-space grid-based Green function calculations. Benchmarks on various NVIDIA GPUs demonstrate the solver’s efficiency, in some cases achieving over 56% of peak floating-point performance for block-sparse matrix multiplications. is open-source, providing interfaces for C, C++, Fortran, Julia, and Python, making it a versatile tool for high-performance computing applications that can benefit from the unification of RHS problems. Paul F. Baumeister, Stepan Nassyr |
J. Supercomput. | 2 |
| 2024 | Exploring Processor Micro-architectures Optimised for BLAS3 Micro-kernels
Stepan Nassyr, Dirk Pleiter |
Euro-Par (2) | 1 |
| 2024 | Application-Driven Exascale: The JUPITER Benchmark SuiteabstractBenchmarks are essential in the design of modern HPC installations, as they define key aspects of system components. Beyond synthetic workloads, it is crucial to include real applications that represent user requirements into benchmark suites, to guarantee high usability and widespread adoption of a new system. Given the significant investments in leadership-class supercomputers of the exascale era, this is even more important and necessitates alignment with a vision of Open Science and reproducibility. In this work, we present the JUPITER Benchmark Suite, which incorporates 16 applications from various domains. It was designed for and used in the procurement of JUPITER, the first European exascale supercomputer. We identify requirements and challenges and outline the project and software infrastructure setup. We provide descriptions and scalability studies of selected applications and a set of key takeaways. The JUPITER Benchmark Suite is released as open source software with this work at github.com/FZJ-JSC/jubench Andreas Herten, Sebastian Achilles, Damian Alvarez, Jayesh Badwaik, Eric Behle, Mathis Bode, Thomas Breuer, Daniel Caviedes-Voullième, Mehdi Cherti, Adel Dabah, Salem El Sayed, Wolfgang Frings, Ana Gonzalez-Nicolas, Eric B. Gregory, Kaveh Haghighi Mood, Thorsten Hater, Jenia Jitsev, Chelsea Maria John, Jan H. Meinke, Catrin I. Meyer, Pavel Mezentsev, Jan-Oliver Mirus, Stepan Nassyr, Carolin Penke, Manoel Römmer, Ujjwal Sinha, Benedikt von St. Vieth, Olaf Stein, Estela Suarez, Dennis Willsch, Ilya Zhukov |
SC | 23 |
| 2020 | Porting Applications to Arm-based ProcessorsabstractArm-based server processors are becoming increasingly used for building massively parallel HPC systems. This triggers the need for porting HPC applications to such architectures and to collect more knowledge about performance benefits and challenges. In this contribution, we report on experiences made during our ongoing efforts to port applications to Arm-based platforms at the Jülich Supercomputing Centre (JSC). The performance of these applications is explored on different Arm-based node architectures plus an x86-based architecture for reference. Bine Brank, Stepan Nassyr, Fatemeh Pouyan, Dirk Pleiter |
CLUSTER | 2 |