EDBT 2026 Demo / reviewers in the wild / expert
Olaf Schenk
dblp:56/4347
· DBLP profile ↗
37ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-8636-1023ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parallel Quadratic Selected Inversion in Quantum Transport SimulationabstractDriven by Moore’s law, the dimensions of transistors have been pushed down to the nanometer scale so that advanced quantum transport (QT) solvers are nowadays required to reliably design such nano-devices. The non-equilibrium Green’s function (NEGF) formalism is suited to this task but is computationally intensive, involving the selected inversion (SI) and the selected solution of quadratic matrix (SQ) equations. Existing algorithms to tackle these numerical problems are ideally suited to GPU acceleration, e.g., the recursive Green’s function (RGF) technique. However, they are typically sequential, limited to block-tridiagonal (BT) matrices, and their implementation has been restricted so far to shared-memory parallelism, limiting the achievable device sizes. To address these shortcomings, we introduce distributed methods that build on RGF and enable parallel SI and SQ. We further extend them to handle BT matrices with arrowhead, allowing for the inclusion of gate leakage currents, a major limiting factor at ultra-scaled device dimensions. We evaluate the performance of our approach on a real dataset from the QT simulation of a nano-ribbon field-effect transistor and perform a comparison with the sparse direct solvers PARDISO and cuDSS. Our SI solver is at least one order of magnitude faster than PARDISO (cuDSS) on CPUs (GPUs), regardless of the system size. When fused, our SI+SQ implementation outperforms the SI-only module of PARDISO by a factor of 1.56 × for the same device dimensions. Performing weak scaling up to 8 CPUs (GPU), our SI+SQ solver achieves a parallel efficiency of \(\eta = 17.1\%\) (\(\eta = 18.5\%\)), thus enabling distributed memory nano-device simulations. Vincent Maillou, Matthias Bollhöfer, Olaf Schenk, Alexandros Nikolaos Ziogas, Mathieu Luisier |
ICS | 3 |
| 2025 | Parallel Selected Inversion of Block-Tridiagonal with Arrowhead MatricesabstractThe inversion of structured sparse matrices is a fundamental yet computationally and memory-intensive task in many scientific applications, such as Bayesian statistical modeling and material science. In certain cases, only particular entries of the full inverse are required. This has motivated the development of so-called selected inversion algorithms (SIA), capable of computing only specific elements of the full inverse. Currently, most SIA implementations are restricted to shared-/distributed-memory CPU architectures or to single GPUs. Here, we introduce novel numerical methods to perform the parallel selected inversion and Cholesky decomposition of positive-definite, block-tridiagonal with arrowhead matrices. A distributed memory, GPU-accelerated implementation of our approach is presented and integrated into the structured solver library Serinv. We demonstrate its performance on synthetic and real datasets from statistical air temperature prediction models and achieve CPU (GPU) speedups of up to$2.6 \times(71.4 \times)$over the SIA of the PARDISO library and up to$14 \times(380.9 \times)$over the MUMPS library, when scaling to 16 processes. Vincent Maillou, Lisa Gaedke-Merzhäuser, Alexandros Nikolaos Ziogas, Olaf Schenk, Mathieu Luisier |
CLUSTER | 4 |
| 2025 | Accelerated Spatio-Temporal Bayesian Modeling for Multivariate Gaussian ProcessesabstractMultivariate Gaussian processes (GPs) offer a powerful probabilistic framework to represent complex interdependent phenomena. They pose, however, significant computational challenges in high-dimensional settings, which frequently arise in spatio-temporal applications. We present DALIA, a highly scalable framework for performing Bayesian inference tasks on spatio-temporal multivariate GPs, based on the methodology of integrated nested Laplace approximations. Our approach relies on a sparse inverse covariance matrix formulation of the GP, puts forward a GPU-accelerated block-dense approach, and introduces a hierarchical, triple-layer, distributed-memory parallel scheme. We showcase weak-scaling performance surpassing the state of the art by two orders of magnitude on a model whose parameter space is 8 × larger and measure strong-scaling speedups of three orders of magnitude when running on 496 GH200 superchips on the Alps supercomputer. Applying DALIA to an air pollution study over northern Italy spanning 48 days, we showcase refined spatial resolutions over the aggregated pollutant measurements. Lisa Gaedke-Merzhäuser, Vincent Maillou, Fernando Rodriguez Avellaneda, Olaf Schenk, Paula Moraga, Mathieu Luisier, Alexandros Nikolaos Ziogas, Håvard Rue |
SC | 4 |
| 2025 | The European master for HPC curriculumabstractInternational audience Pascal Bouvry, Mats Brorsson, Ramon Canal, Aryan Eftekhari, Siegfried Höfinger, Didier Smets, Harald Köstler, Tomás Kozubek, Ezhilmathi Krishnasamy, Josep Llosa, Alexandra Lukas-Rother, Xavier Martorell, Dirk Pleiter, Ana Proykova, Maria-Ribera Sancho, Olaf Schenk, Cristina Silvano |
J. Parallel Distributed Comput. | 16 |
| 2024 | Algorithm 1042: Sparse Precision Matrix Estimation with SQUICabstractWe present SQUIC , a fast and scalable package for sparse precision matrix estimation. The algorithm employs a second-order method to solve the \(\ell_{1}\) -regularized maximum likelihood problem, utilizing highly optimized linear algebra subroutines. In comparative tests using synthetic datasets, we demonstrate that SQUIC not only scales to datasets of up to a million random variables but also consistently delivers runtimes that are significantly faster than other well-established sparse precision matrix estimation packages. Furthermore, we showcase the application of the introduced package in classifying microarray gene expressions. We demonstrate that by utilizing a matrix form of the tuning parameter (also known as the regularization parameter), SQUIC can effectively incorporate prior information into the estimation procedure, resulting in improved application results with minimal computational overhead. Aryan Eftekhari, Lisa Gaedke-Merzhäuser, Dimosthenis Pasadakis, Matthias Bollhöfer, Simon Scheidegger, Olaf Schenk |
ACM Trans. Math. Softw. | 6 |
| 2023 | Sparse Quadratic Approximation for Graph Learningabstract-regularized Gaussian maximum-likelihood method is a popular approach, but also one that poses computational challenges for large scale datasets. Recently proposed methods cast this problem as a constrained optimization variant of precision matrix estimation. In this paper, we build on a state-of-the-art sparse precision matrix estimation method and introduce two algorithms that learn M-matrices, that can be subsequently used for the estimation of graph Laplacian matrices. In the first one, we propose an unconstrained method that follows a post processing approach in order to learn an M-matrix, and in the second one, we implement a constrained approach based on sequential quadratic programming. We also demonstrate the effectiveness, accuracy, and performance of both algorithms. Our numerical examples and comparative results with modern open-source packages reveal that the proposed methods can accelerate the learning of graphs by up to 3 orders of magnitude, while accurately retrieving the latent graphical structure of the data. Furthermore, we conduct large scale case studies for the clustering of COVID-19 daily cases and the classification of image datasets to highlight the applicability in real-world scenarios. Dimosthenis Pasadakis, Matthias Bollhöfer, Olaf Schenk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Level-Based Blocking for Sparse Matrices: Sparse Matrix-Power-Vector MultiplicationabstractThe multiplication of a sparse matrix with a dense vector (SpMV) is a key component in many numerical schemes and its performance is known to be severely limited by main memory access. Several numerical schemes require the multiplication of a sparse matrix polynomial with a dense vector which is typically implemented as a sequence of SpMVs. This results in low performance and ignores the potential to increase the arithmetic intensity by reusing the matrix data from cache. In this work we use the recursive algebraic coloring engine (RACE) to enable blocking of sparse matrix data across the polynomial computations. In the graph representing the sparse matrix we form levels using a breadth-first search. Locality relations of these levels are then used to improve spatial and temporal locality when accessing the matrix data and to implement an efficient multithreaded parallelization. Our approach is independent of the matrix structure and avoids shortcomings of existing “blocking” strategies in terms of hardware efficiency and parallelization overhead. We quantify the quality of our implementation using performance modelling and demonstrate speedups of up to 3× and 5× compared to an optimal SpMV-based baseline on a single multicore chip of recent Intel and AMD architectures. Various numerical schemes like$s$-step Krylov solvers, polynomial preconditioners and power clustering algorithms will benefit from our development. Christie L. Alappat, Georg Hager, Olaf Schenk, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Multiway p-spectral graph cuts on Grassmann manifoldsabstractAbstract Nonlinear reformulations of the spectral clustering method have gained a lot of recent attention due to their increased numerical benefits and their solid mathematical background. We present a novel direct multiway spectral clustering algorithm in thep-norm, for $$p\in (1,2]$$ p∈(1,2] . The problem of computing multiple eigenvectors of the graphp-Laplacian, a nonlinear generalization of the standard graph Laplacian, is recasted as an unconstrained minimization problem on a Grassmann manifold. The value ofpis reduced in a pseudocontinuous manner, promoting sparser solution vectors that correspond to optimal graph cuts aspapproaches one. Monitoring the monotonic decrease of the balanced graph cuts guarantees that we obtain the best available solution from thep-levels considered. We demonstrate the effectiveness and accuracy of our algorithm in various artificial test-cases. Our numerical examples and comparative results with various state-of-the-art clustering methods indicate that the proposed method obtains high quality clusters both in terms of balanced graph cut metrics and in terms of the accuracy of the labelling assignment. Furthermore, we conduct studies for the classification of facial images and handwritten characters to demonstrate the applicability in real-world datasets. Dimosthenis Pasadakis, Christie L. Alappat, Olaf Schenk, Gerhard Wellein |
Mach. Learn. | 3 |
| 2021 | Guest editorial: Virtual special issue on parallel matrix algorithms and applications (PMAA'18)
Olaf Schenk, Peter Arbenz, Luc Giraud, Wim Vanroose |
Parallel Comput. | 1 |
| 2018 | Rethinking large-scale Economic Modeling for Efficiency: Optimizations for GPU and Xeon Phi ClustersabstractWe propose a massively parallelized and optimized framework to solve high-dimensional dynamic stochastic economic models on modern GPU-and KNL-based clusters. First, we introduce a novel approach for adaptive sparse grid index compression alongside a surplus matrix reordering, which significantly reduces the global memory throughput of the compute kernels and maps randomly accessed data onto cache or fast shared memory. Second, we fully vectorize the compute kernels for AVX, AVX2 and AVX512 CPUs, respectively. Third, we develop a hybrid cluster oriented work-preempting scheduler based on TBB, which evenly distributes the time iteration workload onto available CPU cores and accelerators. Numerical experiments on Cray XC40 KNL "Grand Tave" and on Cray XC50 "Piz Daint" systems at the Swiss National Supercomputer Centre (CSCS) show that our framework scales nicely to at least 4,096 compute nodes, resulting in an overall speedup of more than four orders of magnitude compared to a single, optimized CPU thread. As an economic application, we compute global solutions to an annually calibrated stochastic public finance model with sixteen discrete, stochastic states with unprecedented performance. Simon Scheidegger, Dmitry Mikushin, Felix Kubler, Olaf Schenk |
IPDPS | 4 |
| 2018 | Highly Scalable Stencil-Based Matrix-Free Stochastic Estimator for the Diagonal of the InverseabstractSelected inversion problems must be addressed in several research fields like physics, genetics, weather forecasting, and finance, in order to extract selected entries from the inverse of large, sparse matrices. State-of-the-art algorithms are either based on the LU factorization or on an iterative process. Both approaches present computational bottlenecks related to prohibitive memory requirements or extremely high running time for large-scale matrices. In recent years, in order to overcome such limitations, an alternative approach for computing stochastic estimates of the inverse entries has been developed. In this work, we present a stochastic estimator for the diagonal of the inverse and test its performance on a dataset of symmetric, positive semidefinite matrices coming from the field of atomistic quantum transport simulations with nonequilibrium Green's functions (NEGF) formalism. In such a framework, it is required to solve the Schrödinger equation thousands of times, demanding the computation of the diagonal of the retarded Green's function, i.e., the inverse of a large, sparse matrix including open boundary conditions. Given the nature and the structure of the NEGF matrices, our stochastic estimation framework exploits the capabilities of a stencil-based, matrix-free code, avoiding the fill-in and lack of scalability that the LV-based methods present for three-dimensional nanoelectronic devices. We also illustrate the impact of the stochastic estimator by comparing its accuracy against existing methods and demonstrate its scalability performance on the “Piz Daint” cluster at the Swiss National Supercomputing Center, preparing for postpetascale three-dimensional nanoscale calculations. Fabio Verbosio, Jurai Kardos, Mauro Bianco, Olaf Schenk |
SBAC-PAD | 4 |
| 2018 | Multicore Performance Engineering of Sparse Triangular Solves Using a Modified Roofline ModelabstractThe Roofline model is widely used to visualize the performance of executed code together with the upper performance bounds given by the memory bandwidth and the processor peak performance. The model can thus provide an insightful visualization of bottlenecks. In this paper, we try to establish realistic bandwidth ceilings for the sparse triangular solve step of PARDISO, a leading sparse direct solver package, which is also part of the Intel MKL library. The performance of the forward and backward substitution process is analyzed and benchmarked for a representative set of sparse matrices on seven modern x86-type multicore architectures and the Knights Landing manycore architecture. It is shown how to accurately measure the necessary quantities also for threaded code, and the measurement approach, its validation, as well as limitations are discussed. Our modeling approach covers the serial and parallel execution phases, allowing for in-socket performance predictions. Markus Wittmann, Georg Hager, Radim Janalík, Martin Lanser, Axel Klawonn, Oliver Rheinbach, Olaf Schenk, Gerhard Wellein |
SBAC-PAD | 7 |
| 2018 | Distributed memory sparse inverse covariance matrix estimation on high-performance computing architectures
Aryan Eftekhari, Matthias Bollhöfer, Olaf Schenk |
SC | 3 |
| 2018 | Special issue on parallel matrix algorithms and applications (PMAA'16)
Emmanuel Agullo, Peter Arbenz, Luc Giraud, Olaf Schenk |
Parallel Comput. | 4 |
| 2017 | Special issue: Advanced stencil-code engineeringabstractHere, stencil codes are compute-intensive algorithms, in which data points arranged in a large grid are being recomputed repeatedly from the values of data points in a predefined neighborhood. This fixed neighborhood pattern is called a stencil. Stencils codes see wide-spread use in computing the discrete solutions of partial differential equations and systems composed of such equations. Christian Lengauer, Matthias Bolten, Robert D. Falgout, Olaf Schenk |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Special issue on Parallel Matrix Algorithms and Applications (PMAA'14)
Peter Arbenz, Laura Grigori, Rolf Krause, Olaf Schenk |
Parallel Comput. | 4 |
| 2015 | Load-Balanced Local Time Stepping for Large-Scale Wave PropagationabstractIn complex acoustic or elastic media, finite element meshes often require regions of refinement to honour external or internal topography, or small-scale features. These localized smaller elements create a bottleneck for explicit time-stepping schemes due to the Courant-Friedrichs-Lewy stability condition. Recently developed local time stepping (LTS) algorithms reduce the impact of these small elements by locally adapting the time-step size to the size of the element. The recursive, multi-level nature of our LTS scheme introduces an additional challenge, as standard partitioning schemes create a strong load imbalance across processors. We examine the use of multi-constraint graph and hypergraph partitioning tools to achieve effective, load-balanced parallelization. We implement LTS-Newmark in the seismology code SPECFEM3D and compare performance and scalability between different partitioning tools on CPU and GPU clusters using examples from computational seismology. Max Rietmann, Daniel Peter 0001, Olaf Schenk, Bora Uçar, Marcus J. Grote |
IPDPS | 3 |
| 2015 | Towards Parallel Large-Scale Genomic Prediction by Coupling Sparse and Dense Matrix AlgebraabstractGenomic prediction for plant breeding requires taking into account environmental effects and variations of genetic effects across environments. The latter can be modelled by estimating the effect of each genetic marker in every possible environmental condition, which leads to a huge amount of effects to be estimated. Nonetheless, the information about these effects is only sparsely present, due to the fact that plants are only tested in a limited number of environmental conditions. In contrast, the genotypes of the plants are a dense source of information and thus the estimation of both types of effects in one single step would require as well dense as sparse matrix formalisms. This paper presents a way to efficiently apply a high performance computing infrastructure for dealing with large-scale genomic prediction settings, relying on the coupling of dense and sparse matrix algebra. Arne De Coninck, Drosos Kourounis, Fabio Verbosio, Olaf Schenk, Bernard De Baets, Steven Maenhout, Jan Fostier |
PDP | 4 |
| 2015 | Special issue on Parallel Matrix Algorithms and Applications (PMAA'14)
Peter Arbenz, Laura Grigori, Rolf Krause, Olaf Schenk |
Parallel Comput. | 4 |
| 2014 | Fast parallel algorithms for graph similarity and matching
Giorgios Kollias, Madan Sathe, Olaf Schenk, Ananth Grama |
J. Parallel Distributed Comput. | 3 |
| 2014 | Parallel matrix algorithms
Costas Bekas, Ananth Grama, Yousef Saad, Olaf Schenk |
Parallel Comput. | 4 |
| 2013 | Fast Methods for Computing Selected Elements of the Green's Function in Massively Parallel Nanoelectronic Device Simulations
Andrey Kuzmin, Mathieu Luisier, Olaf Schenk |
Euro-Par | 3 |
| 2012 | Patus for convenient high-performance stencils: evaluation in earthquake simulationsabstractPATUS is a code generation and auto-tuning framework for stencil computations targeting modern multi and many-core processors. The goals of the framework are productivity and portability for achieving high performance on the target platform. Its stencil specification language allows the programmer to express the computation in a concise way independently of hardware architecture-specific details. Thus, it increases the programmer productivity by removing the need for manual low-level tuning. We illustrate the impact of the stencil code generation in seismic applications, for which both weak and strong scaling are important. We evaluate the performance by focusing on a scalable discretization of the wave equation and testing complex simulation types of the AWP-ODC code to aim at excellent parallel efficiency, preparing for petascale 3-D earthquake calculations. Matthias Christen, Olaf Schenk, Yifeng Cui |
SC | 2 |
| 2012 | Forward and adjoint simulations of seismic wave propagation on emerging large-scale GPU architecturesabstractComputational seismology is an area of wide sociological and economic impact, ranging from earthquake risk assessment to subsurface imaging and oil and gas exploration. At the core of these simulations is the modeling of wave propagation in a complex medium. Here we report on the extension of the high-order finite-element seismic wave simulation package SPECFEM3D to support the largest scale hybrid and homogeneous supercomputers. Starting from an existing highly tuned MPI code, we migrated to a CUDA version. In order to be of immediate impact to the science mission of computational seismologists, we had to port the entire production package, rather than just individual kernels. One of the challenges in parallelizing finite element codes is the potential for race conditions during the assembly phase. We therefore investigated different methods such as mesh coloring or atomic updates on the GPU. In order to achieve strong scaling, we needed to ensure good overlap of data motion at all levels, including internode and host-accelerator transfers. Finally we carefully tuned the GPU implementation. The new MPI/CUDA solver exhibits excellent scalability and achieves speedup on a node-to-node basis over the carefully tuned equivalent multi-core MPI solver. To demonstrate the performance of both the forward and adjoint functionality, we present two case studies run on the Cray XE6 CPU and Cray XK6 GPU architectures up to 896 nodes: (1) focusing on most commonly used forward simulations, we simulate seismic wave propagation generated by earthquakes in Turkey, and (2) testing the most complex seismic inversion type of the package, we use ambient seismic noise to image 3-D crust and mantle structure beneath western Europe. Max Rietmann, Peter Messmer, Tarje Nissen-Meyer, Daniel Peter 0001, Piero Basini, Dimitri Komatitsch, Olaf Schenk, Jeroen Tromp, Lapo Boschi, Domenico Giardini |
SC | 7 |
| 2012 | An auction-based weighted matching implementation on massively parallel architectures
Madan Sathe, Olaf Schenk, Helmar Burkhart |
Parallel Comput. | 2 |
| 2011 | PATUS: A Code Generation and Autotuning Framework for Parallel Iterative Stencil Computations on Modern MicroarchitecturesabstractStencil calculations comprise an important class of kernels in many scientific computing applications ranging from simple PDE solvers to constituent kernels in multigrid methods as well as image processing applications. In such types of solvers, stencil kernels are often the dominant part of the computation, and an efficient parallel implementation of the kernel is therefore crucial in order to reduce the time to solution. However, in the current complex hardware micro architectures, meticulous architecture-specific tuning is required to elicit the machine's full compute power. We present a code generation and auto-tuning framework \textsc{Patus} for stencil computations targeted at multi- and many core processors, such as multicore CPUs and graphics processing units, which makes it possible to generate compute kernels from a specification of the stencil operation and a parallelization and optimization strategy, and leverages the auto tuning methodology to optimize strategy-dependent parameters for the given hardware architecture. Matthias Christen, Olaf Schenk, Helmar Burkhart |
IPDPS | 2 |
| 2011 | Special issue on Parallel Matrix Algorithms and Applications (PMAA'10)
Peter Arbenz, Yousef Saad, Ahmed H. Sameh, Olaf Schenk |
Parallel Comput. | 4 |
| 2009 | Solving Bi-objective Many-Constraint Bin Packing Problems in Automobile Sheet Metal Forming Processes
Madan Sathe, Olaf Schenk, Helmar Burkhart |
EMO | 2 |
| 2009 | PSPIKE: A Parallel Hybrid Sparse Linear System Solver
Murat Manguoglu, Ahmed H. Sameh, Olaf Schenk |
Euro-Par | 3 |
| 2009 | Parallel data-locality aware stencil computations on modern micro-architecturesabstractNovel micro-architectures including the Cell Broadband Engine Architecture and graphics processing units are attractive platforms for compute-intensive simulations. This paper focuses on stencil computations arising in the context of a biomedical simulation and presents performance benchmarks on both the Cell BE and GPUs and contrasts them with a benchmark on a traditional CPU system. Due to the low arithmetic intensity of stencil computations, typically only a fraction of the peak performance of the compute hardware is reached. An algorithm is presented, which reduces the bandwidth requirements and thereby improves performance by exploiting temporal locality of the data. We report on performance improvements over CPU implementations. Matthias Christen, Olaf Schenk, Esra Neufeld, Peter Messmer, Helmar Burkhart |
IPDPS | 2 |
| 2008 | Algorithmic performance studies on graphics processing units
Olaf Schenk, Matthias Christen, Helmar Burkhart |
J. Parallel Distributed Comput. | 1 |
| 2005 | Special section: SPEEDUP Workshop on Modern algorithms in computational science and information technology
Peter Arbenz, Helmar Burkhart, Erik Maehle, Olaf Schenk |
Future Gener. Comput. Syst. | 4 |
| 2004 | Task-Queue Based Hybrid Parallelism: A Case Study
Karl Fürlinger, Olaf Schenk, Michael Hagemann |
Euro-Par | 2 |
| 2004 | Solving unsymmetric sparse systems of linear equations with PARDISO
Olaf Schenk, Klaus Gärtner |
Future Gener. Comput. Syst. | 1 |
| 2004 | The effects of unsymmetric matrix permutations and scalings in semiconductor device and circuit simulationabstractThe solution of large sparse unsymmetric linear systems is a critical and challenging component of semiconductor device and circuit simulations. The time for a simulation is often dominated by this part. The sparse solver is expected to balance different, and often conflicting requirements. Reliability, a low memory-footprint, and a short solution time are a few of these demands. Currently, no black-box solver exists that can satisfy all criteria. The linear systems from both simulations can be highly ill-conditioned and are, therefore, quite challenging for direct and iterative methods. In this paper, it is shown that algorithms to place large entries on the diagonal using unsymmetric permutations and scalings greatly enhance the reliability of both direct and preconditioned iterative solvers for unsymmetric linear systems arising in semiconductor device and circuit simulations. The numerical experiments indicate that the overall solution strategy is both reliable and cost effective. Olaf Schenk, Stefan Röllin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2002 | Two-level dynamic scheduling in PARDISO: Improved scalability on shared memory multiprocessing systems
Olaf Schenk, Klaus Gärtner |
Parallel Comput. | 1 |
| 2001 | PARDISO: a high-performance serial and parallel sparse linear solver in semiconductor device simulation
Olaf Schenk, Klaus Gärtner, Wolfgang Fichtner, Andreas Stricker |
Future Gener. Comput. Syst. | 1 |