EDBT 2026 Demo / reviewers in the wild / expert
Mark S. Gordon
dblp:76/6591
· DBLP profile ↗
12ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Runtime performance of a GAMESS quantum chemistry application offloaded to GPUsabstractSummary Computational chemistry is at the forefront of solving urgent societal problems, such as polymer upcycling and carbon capture. The complexity of modeling these processes at appropriate length and time scales is mainly manifested in the number and types of chemical species involved in the reactions and may require models of several thousand atoms and large basis sets to accurately capture the chemical complexity and heterogeneity in the physical and chemical processes. The quantum chemistry package General Atomic and Molecular Electronic Structure System (GAMESS) has a wide array of methods that can efficiently and accurately treat complex chemical systems. In this work, we have used the GAMESS Effective Fragment Molecule Orbital (EFMO) method for electronic structure calculation of a challenging mesoporous silica nanoparticle (MSN) model surrounded by about 4700 water molecules to investigate the strong scaling and GPU offloading on hybrid CPU‐GPU nodes. Experiments were performed on the Perlmutter platform at the National Energy Research Scientific Computing Center. Good strong scaling and load balancing have been observed on up to 88 hybrid nodes for different settings of the execution parameters for the calculation considered here. When GPUs are oversubscribed by offloading work from multiple CPU processes, using the NVIDIA multi‐process service (MPS) has consistently reduced time to solution and energy consumed. Additionally, for some configuration parameter settings, oversubscription with MPS improved performance by up to 5.8% over the case without oversubscription. Masha Sosonkina, Gabriel Mateescu, Peng Xu 0054, Tosaporn Sattasathuchana, Buu Pham, Mark S. Gordon, Sarom S. Leang |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on SummitabstractCorrelated electronic structure calculations enable an accurate prediction of the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. The Fragment Molecular Orbital (FMO) method is arguably one of the most effective ways to lower this computational cost while retaining predictive accuracy. In this paper, a novel distributed many-GPU algorithm and implementation of the FMO method are presented. When applied in tandem with the Hartree-Fock and RI-MP2 methods, the new implementation enables correlated calculations on 623,016 electrons and 146,592 atoms in less than 45 minutes using 99.8% of the Summit supercomputer (27,600 GPUs). The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, and excellent strong scalability on Summit achieving 94.6 % parallel efficiency on 4600 nodes. This work makes feasible correlated quantum chemistry calculations on significantly larger molecular systems than before and with higher accuracy. Giuseppe M. J. Barca, Calum Snowdon, Jorge L. Galvez Vallejo, Fazeleh S. Kazemian, Alistair P. Rendell, Mark S. Gordon |
SC | 6 |
| 2021 | Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summitabstractSecond-order Møller-Plesset perturbation theory using the Resolution-of-the-Identity approximation (RI-MP2) is a state-of-the-art approach to accurately estimate many-body electronic correlation effects. This is critical for predicting the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. In this paper, a novel many-GPU algorithm and implementation of a molecular-fragmentation-based RI-MP2 method are presented that enable correlated calculations on over 180,000 electrons and 45,000 atoms using up to the entire Summit supercomputer in 12 minutes. The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, excellent strong scalability on Summit achieving 89.1% parallel efficiency on 4600 nodes, and shows nearly-ideal weak scaling up to 612 nodes. This work makes feasible ab initio correlated quantum chemistry calculations on significantly larger molecular scales than before on both large supercomputing systems and on commodity clusters, with a potential for major impact on progress in chemical, physical, biological and engineering sciences. Giuseppe M. J. Barca, Jorge L. Galvez Vallejo, David Poole 0001, Melisa Alkan, Ryan Stocks, Alistair P. Rendell, Mark S. Gordon |
SC | 7 |
| 2020 | Scaling the hartree-fock matrix build on summitabstractUsage of Graphics Processing Units (GPU) has become strategic for simulating the chemistry of large molecular systems, with the majority of top supercomputers utilizing GPUs as their main source of computational horsepower. In this paper, a new fragmentation-based Hartree-Fock matrix build algorithm designed for scaling on many-GPU architectures is presented. The new algorithm uses a novel dynamic load balancing scheme based on a binned shell-pair container to distribute batches of significant shell quartets with the same code path to different GPUs. This maximizes computational throughput and load balancing, and eliminates GPU thread divergence due to integral screening. Additionally, the code uses a novel Fock digestion algorithm to contract electron repulsion integrals into the Fock matrix, which exploits all forms of permutational symmetry and eliminates thread synchronization requirements. The implementation demonstrates excellent scalability on the Summit computer, achieving good strong scaling performance up to 4096 nodes, and linear weak scaling up to 612 nodes. Giuseppe M. J. Barca, David Poole 0001, Jorge L. Galvez Vallejo, Melisa Alkan, Colleen Bertoni, Alistair P. Rendell, Mark S. Gordon |
SC | 7 |
| 2020 | Runtime power allocation approach for GAMESS hybrid CPU-GPU implementationabstractSummary To improve power consumption of applications at the runtime, modern processors provide frequency scaling capabilities, which along with workload optimization, are also available on GPU accelerators. In this work, a runtime strategy is proposed to distribute a given power allocation among the host components and the GPU according to the current application performance and power usage, such that GPU execution is prioritized over CPU for power allocation to maximize application performance. Next, the strategy is tailored to an application, a quantum‐chemistry package GAMESS for ab initio electronic structure calculations. Specifically, GAMESS hybrid CPU–GPU implementation as provided in the Libcchem library is considered. Experiments, performed on a 28‐core node with a Kepler GPU, resulted in performance gains of up to 50% under the proposed strategy and the largest power allocation considered here as compared with the scenario when this allocation was equally distributed among the computing‐platform components. Vaibhav Sundriyal, Masha Sosonkina, David Poole 0004, Mark S. Gordon |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | An efficient MPI/openMP parallelization of the Hartree-Fock method for the second generation of Intel® Xeon Phi™ processorabstractModern OpenMP threading techniques are used to convert the MPI-only Hartree-Fock code in the GAMESS program to a hybrid MPI/OpenMP algorithm. Two separate implementations that differ by the sharing or replication of key data structures among threads are considered, density and Fock matrices. All implementations are benchmarked on a super-computer of 3,000 Intel® Xeon Phi™ processors. With 64 cores per processor, scaling numbers are reported on up to 192,000 cores. The hybrid MPI/OpenMP implementation reduces the memory footprint by approximately 200 times compared to the legacy code. The MPI/OpenMP code was shown to run up to six times faster than the original for a range of molecular system sizes. Vladimir A. Mironov, Yuri Alexeev, Kristopher Keipert, Michael D'Mello, Alexander A. Moskovsky, Mark S. Gordon |
SC | 6 |
| 2015 | Accelerating Mobile Applications through Flip-Flop ReplicationabstractMobile devices have less computational power and poorer Internet connections than other computers. Computation offload, in which some portions of an application are migrated to a server, has been proposed as one way to remedy this deficiency. Yet, partition-based offload is challenging because it requires applications to accurately predict whether mobile or remote computation will be faster, and it requires that the computation be large enough to overcome the cost of shipping state to and from the server. Further, offload does not currently benefit network-intensive applications. Mark S. Gordon, David Ke Hong, Peter M. Chen, Jason Flinn, Scott A. Mahlke, Z. Morley Mao |
MobiSys | 1 |
| 2012 | COMET: Code Offload by Migrating Execution Transparently
Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Z. Morley Mao, Xu Chen 0028 |
OSDI | 1 |
| 2007 | Integrating Performance Tools with Large-Scale Scientific SoftwareabstractModern performance tools provide methods for easy integration into an application for performance evaluation. For a large-scale scientific software package that has been under development for decades and with developers around the world, several obstacles must be overcome in order to utilize modern performance tools and explore performance bottlenecks. In this paper, we present our experience in integrating performance tools with one popular computational chemistry package. We discuss the difficulties we encountered and the mechanisms developed to integrate performance tools into this code. With performance tools integrated, we show one of the initial performance evaluation results, and discuss what other challenges we are facing to conduct performance evaluation for large-scale scientific packages. Meng-Shiou Wu, Jonathan L. Bentz, Masha Sosonkina, Mark S. Gordon, Ricky A. Kendall |
IPDPS | 5 |
| 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI ModelabstractAn important advance in cluster computing is the evolution from single processor clusters to multi-processor SMP clusters. Due to the increased complexity in the memory model on SMP clusters, new approaches are needed for applications that make use of distributed-memory paradigms. This paper presents new communications software developments that are designed to take advantage of SMP cluster hardware. Although the specific focus is on the central field of computational chemistry and materials science, as embodied in the popular electronic structure package GAMESS (General Atomic and Molecular Electronic Structure System), the impact of these new developments will be far broader in scope. Following a summary of the essential features of the distributed data interface (DDI) in the current implementation of GAMESS, the new developments for SMP clusters are described. The advantages of these new features are illustrated using timing benchmarks on several hardware platforms, using a typical computational chemistry application. Ryan M. Olson, Michael W. Schmidt, Mark S. Gordon, Alistair P. Rendell |
SC | 3 |
| 2002 | Performance and Implementation of Distributed Data CPHF and SCF AlgorithmsabstractThis paper describes a novel distributed data parallel self consistent field (SCF) algorithm and the distributed data coupled perturbed Hartree-Fock (CPHF) step of an analytic Hessian algorithm. The distinguishing features of these algorithms are: (a) columns of density and Fock matrices are distributed among processors, (b) pairwise dynamic load balancing and an efficient static load balancer were developed to achieve a good workload, and (c) network communication time is minimized via careful analysis of data flow in the SCF and CPHF algorithms. By using a shared memory model, novel work load balancers, and improved analytic Hessian steps, we have developed codes that achieve superb performance. The performance of the CPHF code is demonstrated on a large biological system. Yuri Alexeev, Michael W. Schmidt, Theresa L. Windus, Mark S. Gordon, Ricky A. Kendall |
CLUSTER | 4 |
| 2002 | A Distributed Data Implementation of Parallel Full CI ProgramabstractA distributed data parallel full CI program is described The implementation of the FCI algorithm is organized in a combined Cl driven approach With extra computation we were able to avoid redundant communication, and convert the collective communication into more efficient point-to-point communication. The network performance is further optimized by improved DDI library. Examples show very good speedup performance on 16 node PC clusters. The application of the code is also demonstrated. Zhengting Gan, Yuri Alexeev, Ricky A. Kendall, Mark S. Gordon |
CLUSTER | 4 |