EDBT 2026 Demo / reviewers in the wild / expert
Giuseppe M. J. Barca
dblp:284/4837 · also Giuseppe Maria Junior Barca
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0001-5109-4279ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | High-Performance, Accurate Large-Scale Quantum Chemistry Calculations on GPU Supercomputers using Coulomb-Perturbed FragmentationabstractPredicting the chemico-physical properties of large molecular systems is a formidable challenge in chemistry and materials science. Traditional quantum mechanical methods, while accurate, have impractical scaling for large molecules with thousands of atoms, which are crucial in the development of novel therapeutics, catalysts, and nanomaterials. To address this, molecular fragmentation algorithms have been proposed to improve scalability and enable extensive parallelism. In this article, we introduce a significant enhancement to the Fragment Molecular Orbital (FMO) method, termed the Coulomb-Perturbed Fragmentation (CPF) method. CPF incorporates algorithmic improvements and implementation enhancements to optimize performance on heterogeneous computing systems equipped with a large number of GPUs. Key developments include a significant simplification of iteratitve self-consistent field (SCF) algorithm, advanced data management through a one-sided communication model, topology-aware optimizations, and a hybrid communication strategy for intra-group exchanges. Moreover, CPF integrates a distributed dynamic multi-layer load balancing scheme to optimise fragment distribution and workload management across nodes and GPUs. Performance evaluations on a 420-atom benzene molecule system comprising 35 fragments reveal that CPF outperforms existing GPU/CPU-based FMO algorithms in both efficiency and accuracy. When deployed on the Gadi supercomputer, CPF achieves over 97% parallel efficiency on 20 nodes, with scalability maintaining above 98% and 90% efficiency in weak-scaling tests for smaller and larger systems, respectively. Notably, CPF matches or exceeds the computational accuracy of conventional FMO methods, marking a substantial progress in the field of computational chemistry for fragmentation-based large-scale molecular modelling. Fazeleh S. Kazemian, Jorge L. Galvez Vallejo, Giuseppe M. J. Barca |
ICPP | 3 |
| 2024 | Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core SystemsabstractBLAS Level 3 operations are essential for scientific computing, but finding the optimal number of threads for multi-threaded implementations on modern multi-core systems is challenging. We present an extension to the Architecture and Data-Structure Aware Linear Algebra (ADSALA) library that uses machine learning to optimize the runtime of all BLAS Level 3 operations. Our method predicts the best number of threads for each operation based on the matrix dimensions and the system architecture. We test our method on two HPC platforms with Intel and AMD processors, using MKL and BLIS as baseline BLAS implementations. We achieve speedups of 1.5 to 3.0 for all operations, compared to using the maximum number of threads. We also analyze the runtime patterns of different BLAS operations and explain the sources of speedup. Our work shows the effectiveness and generality of the ADSALA approach for optimizing BLAS routines on modern multi-core systems. Yufan Xia, Giuseppe M. J. Barca |
IPDPS | 2 |
| 2024 | Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 PotentialsabstractThe accurate simulation of complex biochemical phenomena has historically been hampered by the computational requirements of high-fidelity molecular-modeling techniques. Quantum mechanical methods, such as ab initio wave-function (WF) theory, deliver the desired accuracy, but have impractical scaling for modeling biosystems with thousands of atoms. Combining molecular fragmentation with MP2 perturbation theory, this study presents an innovative approach that enables biomolecular-scale ab initio molecular dynamics (AIMD) simulations at WF theory level. Leveraging the resolution-of-the-identity approximation for Hartree-Fock and MP2 gradients, our approach eliminates computationally intensive four-center integrals and their gradients, while achieving near-peak performance on modern GPU architectures. The introduction of asynchronous time steps minimizes time step latency, overlapping computational phases and effectively mitigating load imbalances. Utilizing up to $\mathbf{9, 4 0 0}$ nodes of Frontier and achieving $\mathbf{5 9 \%}$ (1006.7 PFLOP/s) of its double-precision floating-point peak, our method enables us to break the million-electron and $1 \mathrm{EFLOP} / \mathrm{s}$ barriers for AIMD simulations with quantum accuracy. Ryan Stocks, Jorge L. Galvez Vallejo, Fiona C. Y. Yu, Calum Snowdon, Elise Palethorpe, Jakub Kurzak, Dmytro Bykov, Giuseppe M. J. Barca |
SC | 8 |
| 2023 | A Machine Learning Approach Towards Runtime Optimisation of Matrix MultiplicationabstractThe GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime.We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB. Yufan Xia, Marco De La Pierre, Amanda S. Barnard, Giuseppe M. J. Barca |
IPDPS | 4 |
| 2023 | AliSim-HPC: parallel sequence simulator for phylogeneticsabstractMOTIVATION: Sequence simulation plays a vital role in phylogenetics with many applications, such as evaluating phylogenetic methods, testing hypotheses, and generating training data for machine-learning applications. We recently introduced a new simulator for multiple sequence alignments called AliSim, which outperformed existing tools. However, with the increasing demands of simulating large data sets, AliSim is still slow due to its sequential implementation; for example, to simulate millions of sequence alignments, AliSim took several days or weeks. Parallelization has been used for many phylogenetic inference methods but not yet for sequence simulation. RESULTS: This paper introduces AliSim-HPC, which, for the first time, employs high-performance computing for phylogenetic simulations. AliSim-HPC parallelizes the simulation process at both multi-core and multi-CPU levels using the OpenMP and message passing interface (MPI) libraries, respectively. AliSim-HPC is highly efficient and scalable, which reduces the runtime to simulate 100 large gap-free alignments (30 000 sequences of one million sites) from over one day to 11 min using 256 CPU cores from a cluster with six computing nodes, a 153-fold speedup. While the OpenMP version can only simulate gap-free alignments, the MPI version supports insertion-deletion models like the sequential AliSim. AVAILABILITY AND IMPLEMENTATION: AliSim-HPC is open-source and available as part of the new IQ-TREE version v2.2.3 at https://github.com/iqtree/iqtree2/releases with a user manual at http://www.iqtree.org/doc/AliSim. Nhan Ly-Trong, Giuseppe M. J. Barca, Bui Quang Minh |
Bioinform. | 2 |
| 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on SummitabstractCorrelated electronic structure calculations enable an accurate prediction of the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. The Fragment Molecular Orbital (FMO) method is arguably one of the most effective ways to lower this computational cost while retaining predictive accuracy. In this paper, a novel distributed many-GPU algorithm and implementation of the FMO method are presented. When applied in tandem with the Hartree-Fock and RI-MP2 methods, the new implementation enables correlated calculations on 623,016 electrons and 146,592 atoms in less than 45 minutes using 99.8% of the Summit supercomputer (27,600 GPUs). The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, and excellent strong scalability on Summit achieving 94.6 % parallel efficiency on 4600 nodes. This work makes feasible correlated quantum chemistry calculations on significantly larger molecular systems than before and with higher accuracy. Giuseppe M. J. Barca, Calum Snowdon, Jorge L. Galvez Vallejo, Fazeleh S. Kazemian, Alistair P. Rendell, Mark S. Gordon |
SC | 1 |
| 2021 | Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summitabstractSecond-order Møller-Plesset perturbation theory using the Resolution-of-the-Identity approximation (RI-MP2) is a state-of-the-art approach to accurately estimate many-body electronic correlation effects. This is critical for predicting the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. In this paper, a novel many-GPU algorithm and implementation of a molecular-fragmentation-based RI-MP2 method are presented that enable correlated calculations on over 180,000 electrons and 45,000 atoms using up to the entire Summit supercomputer in 12 minutes. The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, excellent strong scalability on Summit achieving 89.1% parallel efficiency on 4600 nodes, and shows nearly-ideal weak scaling up to 612 nodes. This work makes feasible ab initio correlated quantum chemistry calculations on significantly larger molecular scales than before on both large supercomputing systems and on commodity clusters, with a potential for major impact on progress in chemical, physical, biological and engineering sciences. Giuseppe M. J. Barca, Jorge L. Galvez Vallejo, David Poole 0001, Melisa Alkan, Ryan Stocks, Alistair P. Rendell, Mark S. Gordon |
SC | 1 |
| 2020 | Scaling the hartree-fock matrix build on summitabstractUsage of Graphics Processing Units (GPU) has become strategic for simulating the chemistry of large molecular systems, with the majority of top supercomputers utilizing GPUs as their main source of computational horsepower. In this paper, a new fragmentation-based Hartree-Fock matrix build algorithm designed for scaling on many-GPU architectures is presented. The new algorithm uses a novel dynamic load balancing scheme based on a binned shell-pair container to distribute batches of significant shell quartets with the same code path to different GPUs. This maximizes computational throughput and load balancing, and eliminates GPU thread divergence due to integral screening. Additionally, the code uses a novel Fock digestion algorithm to contract electron repulsion integrals into the Fock matrix, which exploits all forms of permutational symmetry and eliminates thread synchronization requirements. The implementation demonstrates excellent scalability on the Summit computer, achieving good strong scaling performance up to 4096 nodes, and linear weak scaling up to 612 nodes. Giuseppe M. J. Barca, David Poole 0001, Jorge L. Galvez Vallejo, Melisa Alkan, Colleen Bertoni, Alistair P. Rendell, Mark S. Gordon |
SC | 1 |