Thorsten Kurth

dblp:186/2843 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-0832-6198ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 77% 3D vision · 18% Segmentation and scene understanding · 5%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 52% Processor architecture and microarchitecture · 20% GPUs and heterogeneous computing · 20%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Environmental and earth informatics · 47% Bioinformatics and computational biology · 44% Computational science and engineering · 9%
Databases, data mining, and information retrieval
2 papers
Data mining · 74% Machine learning and data management · 26%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural operator
1.422024
Neural Operators with Localized Integral and Differential Kernels · ICML 2024
Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere · ICML 2023
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Attention on the Sphere · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Attention on the Sphere · NeurIPS 2025
Environmental and earth informatics
climate modeling
0.922023
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere · ICML 2023
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
epistasis
0.812024
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression · SC 2024
Bioinformatics and computational biology › genomics
genome-wide association study
0.812024
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression · SC 2024
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
mixed-precision arithmetic
0.812024
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression · SC 2024
GPUs and heterogeneous computing › GPU computing
tensor cores
0.812024
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression · SC 2024
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator
0.712023
Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere · ICML 2023
Environmental and earth informatics
climate science
0.712023
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
Data mining
dataset construction
0.712023
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
High-performance computing
large-scale training
0.622018
Exascale deep learning for climate analytics · SC 2018
Deep learning at 15PF: supervised and semi-supervised classification for scientific data · SC 2017
High-performance computing › supercomputing
exascale computing
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
High-performance computing
performance optimization at scale
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
High-performance computing
scientific computing systems
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Parallel and multicore computing
distributed deep learning training
0.312017
Deep learning at 15PF: supervised and semi-supervised classification for scientific data · SC 2017
Computational science and engineering › partial differential equation solver
PDE solution operators
0.212024
Neural Operators with Localized Integral and Differential Kernels · ICML 2024
Environmental and earth informatics › climate science
climate data analysis
0.112018
Exascale deep learning for climate analytics · SC 2018

Methods — techniques the papers use, named apart from their topics

fourier neural operator · 2.8mixed-precision linear algebra · 2.3cholesky-based solver · 2.3INT8 tensor cores · 2.3stencil methods · 1.5stochastic regression · 1.3regression baseline · 1.3discrete fourier transform · 1.3autoregressive rollout · 1.3numerical quadrature · 0.9neighborhood attention · 0.9CUDA kernel · 0.9discrete-continuous convolutions · 0.8discrete-continuous convolution · 0.8monte carlo simulation · 0.3lattice QCD · 0.3distributed training · 0.3deep learning · 0.3
YearPublicationVenuePosition
2025 Attention on the Sphere
abstract
We introduce a generalized attention mechanism for spherical domains, enabling Transformer architectures to natively process data defined on the two-dimensional sphere - a critical need in fields such as atmospheric physics, cosmology, and robotics, where preserving spherical symmetries and topology is essential for physical accuracy. By integrating numerical quadrature weights into the attention mechanism, we obtain a geometrically faithful spherical attention that is approximately rotationally equivariant, providing strong inductive biases and leading to better performance than Cartesian approaches. To further enhance both scalability and model performance, we propose neighborhood attention on the sphere, which confines interactions to geodesic neighborhoods. This approach reduces computational complexity and introduces the additional inductive bias for locality, while retaining the symmetry properties of our method. We provide optimized CUDA kernels and memory-efficient implementations to ensure practical applicability. The method is validated on three diverse tasks: simulating shallow water equations on the rotating sphere, spherical image segmentation, and spherical depth estimation. Across all tasks, our spherical Transformers consistently outperform their planar counterparts, highlighting the advantage of geometric priors for learning on spherical domains.
Boris Bonev, Max Rietmann, Andrea Paris, Alberto Carpentieri, Thorsten Kurth
NeurIPS5
2024 Neural Operators with Localized Integral and Differential Kernels
abstract
Neural operators learn mappings between function spaces, which is practical for learning solution operators of PDEs and other scientific modeling applications. Among them, the Fourier neural operator (FNO) is a popular architecture that performs global convolutions in the Fourier space. However, such global operations are often prone to over-smoothing and may fail to capture local details. In contrast, convolutional neural networks (CNN) can capture local features but are limited to training and inference at a single resolution. In this work, we present a principled approach to operator learning that can capture local features under two frameworks by learning differential operators and integral operators with locally supported kernels. Specifically, inspired by stencil methods, we prove that we obtain differential operators under an appropriate scaling of the kernel values of CNNs. To obtain local integral operators, we utilize suitable basis representations for the kernels based on discrete-continuous convolutions. Both these approaches preserve the properties of operator learning and, hence, the ability to predict at any resolution. Adding our layers to FNOs significantly improves their performance, reducing the relative L2-error by 34-72% in our experiments, which include a turbulent 2D Navier-Stokes and the spherical shallow water equations.
Miguel Liu-Schiaffini, Julius Berner, Boris Bonev, Thorsten Kurth, Kamyar Azizzadenesheli, Anima Anandkumar
ICML4
2024 Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
abstract
We exploit the widening margin in tensor-core performance between [FP64/FP32/FP16/INT8,FP64/FP32/FP16/FP8/INT8] on NVIDIA [Ampere,Hopper] GPUs to boost the performance of output accuracy-preserving mixed-precision computation of Genome-Wide Association Studies (GWAS) of 305K patients from the UK BioBank, the largest-ever GWAS cohort studied for genetic epistasis using a multivariate approach. Tile-centric adaptive-precision linear algebraic techniques motivated by reducing data motion gain enhanced significance with low-precision GPU arithmetic. At the core of Kernel Ridge Regression (KRR) techniques for GWAS lie compute-bound cubic-complexity matrix operations that inhibit scaling to aspirational dimensions of the population, genotypes, and phenotypes. We accelerate KRR matrix generation by redesigning the computation for Euclidean distances to engage INT8 tensor cores while exploiting symmetry. We accelerate solution of the regularized KRR systems by deploying a new four-precision Cholesky-based solver, which, at 1.805 mixed-precision ExaOp/s on a nearly full Alps system, outperforms the state-of-the-art CPU-only REGENIE GWAS software by five orders of magnitude.
Hatem Ltaief, Rabab Alomairy, Qinglei Cao, Lotfi Slim, Thorsten Kurth, Benedikt Dorschner, Salim Bougouffa, Rached Abdelkhalak, David E. Keyes
SC6
2023 Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere
abstract
Fourier Neural Operators (FNOs) have proven to be an efficient and effective method for resolution-independent operator learning in a broad variety of application areas across scientific machine learning. A key reason for their success is their ability to accurately model long-range dependencies in spatio-temporal data by learning global convolutions in a computationally efficient manner. To this end, FNOs rely on the discrete Fourier transform (DFT), however, DFTs cause visual and spectral artifacts as well as pronounced dissipation when learning operators in spherical coordinates by incorrectly assuming flat geometry. To overcome this limitation, we generalize FNOs on the sphere, introducing Spherical FNOs (SFNOs) for learning operators on spherical geometries. We apply SFNOs to forecasting atmo- spheric dynamics, and demonstrate stable autoregressive rollouts for a year of simulated time (1,460 steps), while retaining physically plausible dynamics. The SFNO has important implications for machine learning-based simulation of climate dynamics that could eventually help accelerate our response to climate change.
Boris Bonev, Thorsten Kurth, Christian Hundt 0002, Jaideep Pathak, Maximilian Baust, Karthik Kashinath, Anima Anandkumar
ICML2
2023 ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation
abstract
Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore's Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to ML experts because of lack of training data and relevant, easy-to-use workflows. We present ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator's macro-scale physical state.The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res) and code (https://leap-stc.github.io/ClimSim) are released openly to support the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.
Sungduk Yu, Walter M. Hannah, Liran Peng, Zhiyuan Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens, Julius Busecke, Nora Loose, Charles Stern, Tom Beucler, Bryce E. Harrop, Benjamin R. Hillman, Andrea M. Jenney, Savannah L. Ferretti, Nana Liu, Anima Anandkumar, Noah D. Brenowitz, Veronika Eyring, Nicholas Geneva, Pierre Gentine, Stephan Mandt, Jaideep Pathak, Akshay Subramaniam, Carl Vondrick, Rose Yu, Laure Zanna, Ryan Abernathey, Fiaz Ahmed, David C. Bader, Pierre Baldi, Elizabeth A. Barnes, Christopher S. Bretherton, Peter M. Caldwell, Wayne Chuang, Yilun Han, Fernando Iglesias-Suarez, Sanket R. Jantre, Karthik Kashinath, Marat Khairoutdinov, Thorsten Kurth, Nicholas J. Lutsko, Po-Lun Ma, Griffin Mooers, J. David Neelin, David A. Randall, Sara Shamekh, Nathan M. Urban, Janni Yuval, Mike Pritchard
NeurIPS45
2021 IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads
abstract
The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.
Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin
ICPP17
2020 Hierarchical Roofline analysis for GPUs: Accelerating performance optimization for the NERSC-9 Perlmutter system
abstract
Summary The Roofline performance model provides an intuitive and insightful approach to identifying performance bottlenecks and guiding performance optimization. In preparation for the next‐generation supercomputer Perlmutter at NERSC, this paper presents a methodology to construct a hierarchical Roofline on NVIDIA GPUs and extends it to support reduced precision and Tensor Cores. The hierarchical Roofline incorporates L1, L2, device memory, and system memory bandwidths into one single figure, and it offers more profound insights into performance analysis than the traditional DRAM‐only Roofline. We use our Roofline methodology to analyze three proxy applications: GPP from BerkeleyGW, HPGMG from AMReX, and conv2d from TensorFlow. In doing so, we demonstrate the ability of our methodology to readily understand various aspects of performance and performance bottlenecks on NVIDIA GPUs and motivate code optimizations.
Charlene Yang, Thorsten Kurth, Samuel Williams 0001
Concurr. Comput. Pract. Exp.2
2019 Eigensolver performance comparison on Cray XC systems
abstract
Summary Hermitian (symmetric) eigenvalue solvers are the core constituents of electronic structure, quantum‐chemistry, and other HPC applications such as Quantum ESPRESSO, VASP, CP2K, and NWChem to name a few. Our understanding of the performance of symmetric eigenvalue algorithms on various hardware is clearly important to the quantum chemistry or condensed matter physics community but in fact goes beyond that community. For instance, big data analytics is increasingly utilizing eigenvalues solvers, in the study of randomized singular value decomposition (SVD) or principal component analysis (PCA). Noise, vibration, and harshness (NVH) is another field where fast and efficient eigenvalue solvers are required. Most eigenvalue solver packages feature numerous different parameters which can be tuned for performance, eg, the number of nodes, number of total ranks, the decomposition of the matrix, etc. In this paper, we investigate the performance of different packages as well as the influence of these knobs on the solver performance.
Brandon Cook 0001, Thorsten Kurth, Jack Deslippe, Pierre Luc Carrier, Nick Hill, Nathan Wichmann
Concurr. Comput. Pract. Exp.2
2019 TensorFlow at Scale: Performance and productivity analysis of distributed training with Horovod, MLSL, and Cray PE ML
abstract
Summary Deep learning has proven to be a successful tool for solving a large variety of problems in various scientific fields and beyond. In recent years, the models as well as the available datasets have grown bigger and more complicated, and thus, an increasing amount of computing resources is required in order to train these models in a reasonable amount of time. Besides being able to use HPC resources, deep learning model developers want flexible frameworks which allow for rapid prototyping. One of the most important of these frameworks is Google TensorFlow, which provides both features, ie, good performance as well as flexibility. In this paper, we discuss different solutions for scaling the TensorFlow Framework to thousands of nodes on contemporary Cray XC supercomputing systems.
Thorsten Kurth, Mikhail Smorkalov, Peter Mendygral, Srinivas Sridharan 0002, Amrita Mathuriya
Concurr. Comput. Pract. Exp.1
2018 Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing
Evan Berkowitz, Michael A. Clark, Arjun Singh Gambhir, Kenneth S. McElvain, Amy N. Nicholson, Enrico Rinaldi, Pavlos Vranas, André Walker-Loud, Chia-Cheng Chang, Bálint Joó, Thorsten Kurth, Konstantinos Orginos
SC11
2018 Exascale deep learning for climate analytics
Thorsten Kurth, Sean Treichler, Joshua Romero, Mayur Mudigonda, Nathan Luehr, Everett H. Phillips, Ankur Mahesh, Michael A. Matheson, Jack Deslippe, Massimiliano Fatica, Prabhat, Michael Houston
SC1
2018 Preparing NERSC users for Cori, a Cray XC40 system with Intel many integrated cores
abstract
Summary The newest NERSC supercomputer Cori is a Cray XC40 system consisting of 2,388 Intel Xeon Haswell nodes and 9,688 Intel Xeon‐Phi “Knights Landing” (KNL) nodes. Compared to the Xeon‐based clusters NERSC users are familiar with, optimal performance on Cori requires consideration of KNL mode settings; process, thread, and memory affinity; fine‐grain parallelization; vectorization; and use of the high‐bandwidth MCDRAM memory. This paper describes our efforts preparing NERSC users for KNL through the NERSC Exascale Science Application Program, Web documentation, and user training. We discuss how we configured the Cori system for usability and productivity, addressing programming concerns, batch system configurations, and default KNL cluster and memory modes. System usage data, job completion analysis, programming and running jobs issues, and a few successful user stories on KNL are presented.
Yun (Helen) He, Brandon Cook 0001, Jack Deslippe, Brian Friesen, Richard A. Gerber, Rebecca Hartman-Baker, Alice E. Koniges, Thorsten Kurth, Stephen Leak, Woo-Sun Yang, Zhengji Zhao, Eddie Baron, Peter Hauschildt
Concurr. Comput. Pract. Exp.8
2017 Deep learning at 15PF: supervised and semi-supervised classification for scientific data
abstract
This paper presents the first, 15-PetaFLOP Deep Learning system for solving scientific pattern classification problems on contemporary HPC architectures. We develop supervised convolutional architectures for discriminating signals in high-energy physics data as well as semi-supervised architectures for localizing and classifying extreme weather in climate data. Our Intelcaffe-based implementation obtains ~2TFLOP/s on a single Cori Phase-II Xeon-Phi node. We use a hybrid strategy employing synchronous node-groups, while using asynchronous communication across groups. We use this strategy to scale training of a single model to ~9600 Xeon-Phi nodes; obtaining peak performance of 11.73-15.07 PFLOP/s and sustained performance of 11.41-13.27 PFLOP/s. At scale, our HEP architecture produces state-of-the-art classification accuracy on a dataset with 10M images, exceeding that achieved by selections on high-level physics-motivated features. Our semi-supervised architecture successfully extracts weather patterns in a 15TB climate dataset. Our results demonstrate that Deep Learning can be optimized and scaled effectively on many-core, HPC systems.
Thorsten Kurth, Jian Zhang 0049, Nadathur Satish, Evan Racah, Ioannis Mitliagkas, Md. Mostofa Ali Patwary, Tareq M. Malas, Narayanan Sundaram, Wahid Bhimji, Mikhail Smorkalov, Jack Deslippe, Mikhail Shiryaev, Srinivas Sridharan 0002, Prabhat, Pradeep Dubey
SC1
2016 MPI usage at NERSC: Present and Future
abstract
In this poster, we describe how MPI is used at the National Energy Research Scientific Computing Center (NERSC) NERSC is the production high-performance computing center for the US Department of Energy, with more than 5000 users and 800 distinct projects. Through a variety of tools (e.g., User Survey, application team collaborations, etc.), we determine how MPI is used on our latest systems, with a particular focus on advanced features and how early applications intend to use MPI on NERSC's upcoming Intel Knights Landing (KNL) many-core system1 - one of the first to be deployed. In the poster, we also compare the usage of MPI to exascale developmental programming models such as UPC++ and HPX, with an eye on what features and extensions to MPI are plausible and useful for NERSC users. We also discuss perceived shortcomings of MPI, and why certain groups use other parallel programming models on the systems. In addition to a broad survey of the NERSC HPC population, we follow the evolution of a few key application codes2 that are being highly optimized for the KNL architecture using advanced OpenMP techniques. We study how these highly optimized on-node proxy apps and full applications start to make the transition to using full hybrid MPI+OpenMP implementations on the self-hosted KNL system.
Alice E. Koniges, Brandon Cook 0001, Jack Deslippe, Thorsten Kurth, Hongzhang Shan
EuroMPI4