VLDB 2026 Research / reviewers in the wild / expert
Carl Pearson
dblp:149/5509
· DBLP profile ↗
4ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0001-6481-970XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at ScaleabstractTrilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors. Matthias Mayr, Alexander Heinlein, Christian A. Glusa, Sivasankaran Rajamanickam, Maarten Arnst, Roscoe A. Bartlett, Luc Berger-Vergiat, Erik G. Bowman, Karen D. Devine, Graham Harper, Michael A. Heroux, Mark Hoemmen, Jonathan J. Hu, Brian Michael Kelley, Kyungjoo Kim, Drew P. Kouri, Paul Kuberry, Kim Liegeois, Curtis C. Ober, Roger P. Pawlowski, Carl Pearson, Mauro Perego, Eric T. Phipps, Denis Ridzal, Nathan V. Roberts, Christopher M. Siefert, Heidi Thornquist, Romin Tomasetti, Christian Trott, Ray S. Tuminaro, James M. Willenbring, Michael M. Wolf, Ichitaro Yamazaki |
ACM Trans. Math. Softw. | 21 |
| 2021 | TEMPI: An Interposed MPI Library with a Canonical Representation of CUDA-aware DatatypesabstractMPI derived datatypes are an abstraction that simplifies handling of non-contiguous data in MPI applications. These datatypes are recursively constructed at runtime from primitive Named Types defined in the MPI standard. More recently, the development and deployment of CUDA-aware MPI implementations has encouraged the transition of distributed high-performance MPI codes to use GPUs. Such implementations allow MPI functions to directly operate on GPU buffers, easing integration of GPU compute into MPI codes. This work first presents a novel datatype handling strategy for nested strided datatypes, which finds a middle ground between the specialized or generic handling in prior work. This work also shows that the performance characteristics of non-contiguous data handling can be modeled with empirical system measurements, and used to transparently improve MPI_Send/Recv latency. Finally, despite substantial attention to non-contiguous GPU data and CUDA-aware MPI implementations, good performance cannot be taken for granted. This work demonstrates its contributions through an MPI interposer library, TEMPI. TEMPI can be used with existing MPI deployments without system or application changes. Ultimately, the interposed-library model of this work demonstrates MPI_Pack speedup of up to 242000x and MPI_Send speedup of up to 59000x compared to the MPI implementation deployed on a leadership-class supercomputer. This yields speedup of more than 917x in a 3D halo exchange with 3072 processes. Carl Pearson, Kun Wu 0002, I-Hsin Chung, Jinjun Xiong, Wen-Mei W. Hwu |
HPDC | 1 |
| 2019 | Evaluating Characteristics of CUDA Communication Primitives on High-Bandwidth InterconnectsabstractData-intensive applications such as machine learning and analytics have created a demand for faster interconnects to avert the memory bandwidth wall and allow GPUs to be effectively leveraged for lower compute intensity tasks. This has resulted in wide adoption of heterogeneous systems with varying underlying interconnects, and has delegated the task of understanding and copying data to the system or application developer. No longer is a malloc followed by memcpy the only or dominating modality of data transfer; application developers are faced with additional options such as unified memory and zero-copy memory. Data transfer performance on these systems is now impacted by many factors including data transfer modality, system interconnect hardware details, CPU caching state, CPU power management state, driver policies, virtual memory paging efficiency, and data placement. Carl Pearson, Abdul Dakkak, Sarah Hashash, Cheng Li 0014, I-Hsin Chung, Jinjun Xiong, Wen-Mei W. Hwu |
ICPE | 1 |
| 2018 | A Fast and Massively-Parallel Inverse Solver for Multiple-Scattering Tomographic Image ReconstructionabstractWe present a massively-parallel solver for large Helmholtz-type inverse scattering problems. The solver employs the distorted Born iterative method for capturing the multiple-scattering phenomena in image reconstructions. This method requires many full-wave forward-scattering solutions in each iteration, constituting the main performance bottleneck with its high computational complexity. As a remedy, we use the multilevel fast multipole algorithm (MLFMA). The solver scales among computing nodes using a two-dimensional parallelization strategy that distributes illuminations in one dimension, and MLFMA sub-trees in the other dimension. Multi-core CPUs and GPUs are used to provide per-node speedup. We demonstrate a 76% efficiency when scaling from 64 GPUs to 4,096 GPUs. The paper provides reconstruction of a 204.8λ×204.8λ image (4M unknowns) executed on 4,096 GPUs in near-real time (almost 2 minutes). To the best of our knowledge, this is the largest full-wave inverse scattering solution to date, in terms of both image size and computational resources. Mert Hidayetoglu, Carl Pearson, Izzat El Hajj, Levent Gürel, Weng Cho Chew, Wen-Mei W. Hwu |
IPDPS | 2 |