Melisa Alkan

dblp:284/5135 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-3110-2945ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 61% GPUs and heterogeneous computing · 30% Parallel and multicore computing · 9%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
quantum chemistry simulation
0.922021
Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021
Scaling the hartree-fock matrix build on summit · SC 2020
High-performance computing
scientific computing systems
0.922021
Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021
Scaling the hartree-fock matrix build on summit · SC 2020
GPUs and heterogeneous computing › multi-GPU computing
distributed GPU computing
0.512021
Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021
GPUs and heterogeneous computing › GPU resource management
GPU load balancing
0.412020
Scaling the hartree-fock matrix build on summit · SC 2020
Parallel and multicore computing › load balancing
dynamic load balancing
0.112020
Scaling the hartree-fock matrix build on summit · SC 2020

Methods — techniques the papers use, named apart from their topics

molecular fragmentation · 0.5many-GPU algorithm · 0.5RI-MP2 · 0.5fragmentation-based algorithm · 0.4fock digestion · 0.4dynamic load balancing · 0.4
YearPublicationVenuePosition
2021 Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit
abstract
Second-order Møller-Plesset perturbation theory using the Resolution-of-the-Identity approximation (RI-MP2) is a state-of-the-art approach to accurately estimate many-body electronic correlation effects. This is critical for predicting the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. In this paper, a novel many-GPU algorithm and implementation of a molecular-fragmentation-based RI-MP2 method are presented that enable correlated calculations on over 180,000 electrons and 45,000 atoms using up to the entire Summit supercomputer in 12 minutes. The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, excellent strong scalability on Summit achieving 89.1% parallel efficiency on 4600 nodes, and shows nearly-ideal weak scaling up to 612 nodes. This work makes feasible ab initio correlated quantum chemistry calculations on significantly larger molecular scales than before on both large supercomputing systems and on commodity clusters, with a potential for major impact on progress in chemical, physical, biological and engineering sciences.
Giuseppe M. J. Barca, Jorge L. Galvez Vallejo, David Poole 0001, Melisa Alkan, Ryan Stocks, Alistair P. Rendell, Mark S. Gordon
SC4
2020 Scaling the hartree-fock matrix build on summit
abstract
Usage of Graphics Processing Units (GPU) has become strategic for simulating the chemistry of large molecular systems, with the majority of top supercomputers utilizing GPUs as their main source of computational horsepower. In this paper, a new fragmentation-based Hartree-Fock matrix build algorithm designed for scaling on many-GPU architectures is presented. The new algorithm uses a novel dynamic load balancing scheme based on a binned shell-pair container to distribute batches of significant shell quartets with the same code path to different GPUs. This maximizes computational throughput and load balancing, and eliminates GPU thread divergence due to integral screening. Additionally, the code uses a novel Fock digestion algorithm to contract electron repulsion integrals into the Fock matrix, which exploits all forms of permutational symmetry and eliminates thread synchronization requirements. The implementation demonstrates excellent scalability on the Summit computer, achieving good strong scaling performance up to 4096 nodes, and linear weak scaling up to 612 nodes.
Giuseppe M. J. Barca, David Poole 0001, Jorge L. Galvez Vallejo, Melisa Alkan, Colleen Bertoni, Alistair P. Rendell, Mark S. Gordon
SC4