Alessandro Curioni

dblp:14/6780 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 97% Parallel and multicore computing · 3%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.422015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
11 PFLOP/s simulations of cloud cavitation collapse · SC 2013
Computational science and engineering
computational chemistry
0.212016
Enhanced MPSM3 for applications to quantum biological simulations · SC 2016
High-performance computing › sparse linear algebra
sparse matrix computation
0.212016
Enhanced MPSM3 for applications to quantum biological simulations · SC 2016
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.212016
Enhanced MPSM3 for applications to quantum biological simulations · SC 2016
High-performance computing › performance optimization at scale
extreme-scale scalability
0.212015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
High-performance computing
performance optimization at scale
0.212015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
High-performance computing › scientific computing systems
computational fluid dynamics
0.212013
11 PFLOP/s simulations of cloud cavitation collapse · SC 2013
High-performance computing › large-scale simulation
extreme-scale simulation
0.212013
11 PFLOP/s simulations of cloud cavitation collapse · SC 2013
High-performance computing › supercomputing
petascale computing
0.212013
11 PFLOP/s simulations of cloud cavitation collapse · SC 2013
Mathematical optimization › numerical analysis › multigrid methods
algebraic multigrid
0.112015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
Mathematical optimization › numerical analysis
multigrid methods
0.112015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015

Methods — techniques the papers use, named apart from their topics

self-consistent field method · 0.5density matrix multiplication · 0.5schur-complement preconditioning · 0.4multi-octree adaptivity · 0.4mixed continuous-discontinuous discretization · 0.4two-phase flow simulation · 0.2performance optimization · 0.2
YearPublicationVenuePosition
2018 A scalable iterative dense linear system solver for multiple right-hand sides in data analytics
Vassilis Kalantzis, Cristiano Malossi, Costas Bekas, Alessandro Curioni, Efstratios Gallopoulos, Yousef Saad
Parallel Comput.4
2016 Key/Value-Enabled Flash Memory for Complex Scientific Workflows with On-Line Analysis and Visualization
abstract
Scientific workflows are often composed of compute-intensive simulations and data-intensive analysis and visualization, both equally important for productivity. High-performance computers run the compute-intensive phases efficiently, but data-intensive processing is still getting less attention. Dense non-volatile memory integrated into super-computers can help address this problem. In addition to density, it offers significantly finer-grained I/O than disk-based I/O systems. We present a way to exploit the fundamental capabilities of Storage-Class Memories (SCM), such as Flash, by using scalable key-value (KV) I/O methods instead of traditional file I/O calls commonly used in HPC systems. Our objective is to enable higher performance for on-line and near-line storage for analysis and visualization of very high resolution, but correspondingly transient, simulation results. In this paper, we describe 1) the adaptation of a scalable key-value store to a BlueGene/Q system with integrated Flash memory, 2) a novel key-value aggregation module which implements coalesced, function-shipped calls between the clients and the servers, and 3) the refactoring of a scientific workflow to use application-relevant keys for fine-grained data subsets. The resulting implementation is analogous to function-shipping of POSIX I/O calls but shows an order of magnitude increase in read and a factor 2.5x increase in write IOPS performance (11 million read IOPS, 2.5 million write IOPS from 4096 compute nodes) when compared to a classical file system on the same system. It represents an innovative approach for the integration of SCM within an HPC system at scale.
Stefan Eilemann, Fabien Delalondre, Jon Bernard, Judit Planas, Felix Schürmann, John Biddiscombe, Costas Bekas, Alessandro Curioni, Bernard Metzler, Peter Kaltstein, Peter Morjan, Joachim Fenkes, Ralph Bellofatto, Lars Schneidenbach, T. J. Christopher Ward, Blake G. Fitch
IPDPS8
2016 Stochastic Matrix-Function Estimators: Scalable Big-Data Kernels with High Performance
abstract
In this era of Big Data, large graphs appear in many scientific domains. To extract the hidden knowledge/correlations in these graphs, novel methods need to be developed to analyse these graphs fast. In this paper, we present a unified framework of stochastic matrix-function estimators, which allows one to compute a subset of elements of the matrix f(A), where f is an arbitrary function and A is the adjacency matrix of the graph. The new framework has a computational cost proportional to the size of the subset, i.e. to obtain the diagonal of f(A) with matrix-size N, the computational cost is proportional to N contrary to the traditional N^3 from diagonalization. Furthermore, we will show that the new framework allows us to write implementations of the algorithm that scale naturally with the number of compute nodes and is easily ported to accelerators where the kernels perform very well.
Peter W. J. Staar, Panagiotis Kl. Barkoutsos, Roxana Istrate, Cristiano Malossi, Ivano Tavernelli, Nikolaj Moll, Heiner Giefers, Christoph Hagleitner, Costas Bekas, Alessandro Curioni
IPDPS10
2016 Enhanced MPSM3 for applications to quantum biological simulations
abstract
Classical molecular dynamics simulations have been the preferred method to cope with the characteristic sizes and time scales of complex life-science systems. However, while classical methods have well known limitations, such as that their accuracy strongly depends on empirical tuning, the practical use of far more accurate methods that rely on quantum Hamiltonians, has been limited by the current efficiency and scalability of sparse matrix-matrix multiplication algorithms used in the self-consistent field equations. In this paper, we show unprecedented massive scalability of a recently presented method, called MPSM3, for sparse matrix-matrix multiplication. The algorithmic basis of the method was presented in a recent publication, while here we describe the algorithmic enhancements that allow us to claim at least one order of magnitude improvement in scalability and time to solution over the state of the art (original MPSM3). We achieve a time to solution for the multiplication of density matrices within the self-consistent field scheme that is approaching the time needed to evaluate energy and forces with classical force-field methods and that is independent from the system size, provided proportional resources. This latest development renders the application of entirely quantum Hamiltonians to systems of several millions of atoms for extended molecular dynamics investigations feasible.
A. Pozdneev, Valéry Weber, Teodoro Laino, Costas Bekas, Alessandro Curioni
SC5
2015 An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle
abstract
Mantle convection is the fundamental physical process within earth's interior responsible for the thermal and geological evolution of the planet, including plate tectonics. The mantle is modeled as a viscous, incompressible, non-Newtonian fluid. The wide range of spatial scales, extreme variability and anisotropy in material properties, and severely nonlinear rheology have made global mantle convection modeling with realistic parameters prohibitive. Here we present a new implicit solver that exhibits optimal algorithmic performance and is capable of extreme scaling for hard PDE problems, such as mantle convection. To maximize accuracy and minimize runtime, the solver incorporates a number of advances, including aggressive multi-octree adaptivity, mixed continuous-discontinuous discretization, arbitrarily-high-order accuracy, hybrid spectral/geometric/algebraic multigrid, and novel Schur-complement preconditioning. These features present enormous challenges for extreme scalability. We demonstrate that---contrary to conventional wisdom---algorithmically optimal implicit solvers can be designed that scale out to 1.5 million cores for severely nonlinear, ill-conditioned, heterogeneous, and anisotropic PDEs.
Johann Rudi, Cristiano Malossi, Tobin Isaac, Georg Stadler, Michael Gurnis, Peter W. J. Staar, Yves Ineichen, Costas Bekas, Alessandro Curioni, Omar Ghattas
SC9
2014 Shedding Light on Lithium/Air Batteries Using Millions of Threads on the BG/Q Supercomputer
abstract
In this work, we present a novel parallelization scheme for a highly efficient evaluation of the Hartree-Fock exact exchange (HFX) in ab initio molecular dynamics simulations, specifically tailored for condensed phase simulations. Our developments allow one to achieve the necessary accuracy for the evaluation of the HFX in a highly controllable manner. We show here that our solutions can take great advantage of the latest trends in HPC platforms, such as extreme threading, short vector instructions and highly dimensional interconnection networks. Indeed, all these trends are evident in the IBM Blue Gene/Q supercomputer. We demonstrate an unprecedented scalability up to 6,291,456 threads (96 BG/Q racks) with a near perfect parallel efficiency, which represents a more than 20-fold improvement as compared to the current state of the art. In terms of reduction of time to solution, we achieved an improvement that can surpass a 10-fold decrease in runtime with respect to directly comparable approaches. We exploit this development to enhance the accuracy of DFT based molecular dynamics by using the PBE0 hybrid functional. This approach allowed us to investigate the chemical behavior of organic solvents in one of the most challenging research topics in energy storage, lithium/air batteries, and to propose alternative solvents with enhanced stability to ensure an appropriate reversible electrochemical reaction. This step is key for the development of a viable lithium/air storage technology, which would have been a daunting computational task using standard methods. Recent research has shown that the electrolyte plays a key role in non-aqueous lithium/air batteries in producing the appropriate reversible electrochemical reduction. In particular, the chemical degradation of propylene carbonate, the typical electrolyte used, by lithium peroxide has been demonstrated by molecular dynamics simulations of highly realistic models. Reaching the necessary high accuracy in these simulations is a daunting computational task using standard methods.
Valéry Weber, Costas Bekas, Teodoro Laino, Alessandro Curioni, Adam Bertsch, Scott Futral
IPDPS4
2013 11 PFLOP/s simulations of cloud cavitation collapse
abstract
We present unprecedented, high throughput simulations of cloud cavitation collapse on 1.6 million cores of Sequoia reaching 55% of its nominal peak performance, corresponding to 11 PFLOP/s. The destructive power of cavitation reduces the lifetime of energy critical systems such as internal combustion engines and hydraulic turbines, yet it has been harnessed for water purification and kidney lithotripsy. The present two-phase flow simulations enable the quantitative prediction of cavitation using 13 trillion grid points to resolve the collapse of 15'000 bubbles. We advance by one order of magnitude the current state-of-the-art in terms of time to solution, and by two orders the geometrical complexity of the flow. The software successfully addresses the challenges that hinder the effective solution of complex flows on contemporary supercomputers, such as limited memory bandwidth, I/O bandwidth and storage capacity. The present work redefines the frontier of high performance computing for fluid dynamics simulations.
Diego Rossinelli, Babak Hejazialhosseini, Panagiotis Hadjidoukas, Costas Bekas, Alessandro Curioni, Adam Bertsch, Scott Futral, Steffen J. Schmidt, Nikolaus A. Adams, Petros Koumoutsakos
SC5
2012 Low-cost data uncertainty quantification
abstract
SUMMARY The analysis of a huge backload of ever‐accumulating data presents a huge challenge in all respects of computing. Inverse covariance matrices in this respect are very important. We target data uncertainty quantification, a very useful measure of which is provided by inverse covariance matrix diagonal entries. In previous work, we introduced a novel method that reduces overall complexity by at least two orders of magnitude. At the same time, a state‐of‐the‐art message‐passing interface (MPI) implementation allowed us to reach a sustained performance of up to 73% (730 TFLOPS on the full 72 Blue Gene/P rack configuration at Jülich). Thanks to its reduced complexity, this work has attracted significant interest, and thus, we have received numerous requests concerning its exploitation in various fields. A common denominator in these requests is that they almost all came from people with no or, in the best case, limited high‐performance computing background. Nevertheless, all interest is in analyzing huge data sets, suitably adapting the method to particular applications. A bottleneck then is that potential users are reluctant to pay for a steep learning curve to get proficient in parallel computing using the de facto standard: MPI. Thus, we turned to the Partitioned Global Address Space programming model and in particular the Unified Parallel C language. In this work, we gave a comprehensive description of the framework and demonstrated the efficiency of the state‐of‐the‐art MPI implementation. In addition, we showed that one can develop an easy‐to‐follow yet efficient Unified Parallel C implementation, which is also easy to debug and maintain, features that significantly boost overall productivity. Copyright © 2011 John Wiley & Sons, Ltd.
Costas Bekas, Alessandro Curioni, Irina Fedulova
Concurr. Comput. Pract. Exp.2
2010 Extreme scalability challenges in micro-finite element simulations of human bone
abstract
Abstract Coupling recent imaging capabilities with microstructural finite element (µFE) analysis offers a powerful tool to determine bone stiffness and strength. It shows high potential to improve the individual fracture risk prediction, a tool much needed in the diagnosis and treatment of osteoporosis, that is, according to the World Health Organization (WHO), second only to cardiovascular disease as a leading health‐care problem. We adapted a multilevel preconditioned conjugate gradient method to solve the very large voxel models that arise in the µFE bone structure analysis. The intricate microstructure properties of bone lead to sparse matrices with billions of rows, thus rendering this application to be an ideal candidate for massively parallel architectures such as the BG/L Supercomputer. In this work we present our progress as well as the challenges we were able to identify in our quest to achieve scalability to thousands of BG/L cores. Copyright © 2010 John Wiley & Sons, Ltd.
Costas Bekas, Alessandro Curioni, Peter Arbenz, Cyril Flaig, G. Harry van Lenthe, Ralph Müller, Andreas J. Wirth
Concurr. Comput. Pract. Exp.2
2008 Atomic wavefunction initialization in ab initio
Costas Bekas, Alessandro Curioni, Wanda Andreoni
Parallel Comput.2
2005 Early Experience with Scientific Applications on the Blue Gene/L Supercomputer
Gheorghe Almási 0001, Gyan Bhanot, Dong Chen 0005, Maria Eleftheriou, Blake G. Fitch, Alan Gara, Robert S. Germain, John A. Gunnels, Manish Gupta 0002, Philip Heidelberger, Michael Pitman, Aleksandr Rayshubskiy, James C. Sexton, Frank Suits, Pavlos Vranas, Robert Walkup, T. J. Christopher Ward, Yuriy Zhestkov, Alessandro Curioni, Wanda Andreoni, Charles Archer, José E. Moreira, Richard Loft, Henry M. Tufo, Theron Voran, Katherine Riley
Euro-Par19
2005 HPCx: towards capability computing
abstract
Abstract We introduce HPCx—the U.K.'s new National HPC Service—which aims to deliver a world‐class service for capability computing to the U.K. scientific community. HPCx is targeting an environment that will both result in world‐leading science and address the challenges involved in scaling existing codes to the capability levels required. Close working relationships with scientific consortia and user groups throughout the research process will be a central feature of the service. A significant number of key user applications have already been ported to the system. We present initial benchmark results from this process and discuss the optimization of the codes and the performance levels achieved on HPCx in comparison with other systems. We find a range of performance with some algorithms scaling far better than others. Copyright © 2005 John Wiley & Sons, Ltd.
Mike Ashworth, Ian J. Bush, Martyn F. Guest, Andrew G. Sunderland, Stephen Booth, Joachim Hein, Lorna Smith, Kevin Stratford, Alessandro Curioni
Concurr. Pract. Exp.9
2005 Dual-level parallelism for ab initio molecular dynamics: Reaching teraflop performance with the CPMD code
Jürg Hutter, Alessandro Curioni
Parallel Comput.2
2000 New advances in chemistry and materials science with CPMD and parallel computing
Wanda Andreoni, Alessandro Curioni
Parallel Comput.2