Nigel Tan

dblp:286/5486 · also Nigel Phillip Tan · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4699-8657ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Modernizing VPIC-Kokkos I/O: From Legacy Binary Output to Adaptive HDF5 Workflows
abstract
Exascale particle-in-cell (PIC) simulations like Vector Particle-In-Cell (VPIC) face critical parallel input/output (I/O) bottlenecks due to legacy proprietary formats and storage bloat from redundant ghost cells. We present two architectural contributions: a Kokkos-aware staging pipeline that eliminates ghost cell padding – yielding a 67% reduction in grid-based export sizes – and a parallel Hierarchical Data Format 5 (HDF5) backend validated against h5bench in a weak-scaling study up to 896 MPI ranks. Our pipelined architecture achieves file-per-process (FPP) throughput comparable to the legacy binary format, outpacing monolithic bulk-writing benchmarks while isolating collective I/O (CIO) synchronization overheads, establishing a definitive performance baseline for emerging Exascale architectures.
Connor Browne, Nigel Tan, Scott V. Luedtke, Michela Taufer, Brian J. Albright
HPDC2
2026 Merkle-Tree Weight Snapshot Deduplication for Provenance-Aware Auditing of Neural Network Training
abstract
Weight snapshots taken during neural network training provide a foundation for reproducibility and for understanding how models evolve during learning. They indicate whether networks progress toward higher accuracy or diverge toward poor generalization, yet their size and frequency impose severe storage and I/O burdens. As models scale, snapshots exhibit substantial cross-epoch redundancy, making them increasingly difficult to archive and analyze efficiently. We introduce a Merkle-tree deduplication pipeline that removes redundancy while exposing metadata about training dynamics. Chunking and deduplicating weights yields 70–80% storage savings across CIFAR-10/100 and protein diffraction datasets, outperforming list-based deduplication and per-snapshot compression baselines. Beyond space savings, Merkle-tree metadata categorizes chunks as fixed duplicates, shifted duplicates, or first occurrences. These signals predict validation accuracy with mean absolute error below 1% and provide an optional, metadata-driven signal to inform early stopping, enabling savings of 16–72% of the training epochs with negligible accuracy loss. Our work demonstrates that Merkle-tree deduplication provides a unified approach to reduce overhead, preserve reproducibility, and explain training dynamics within user-defined error tolerances, without disrupting the learning loop.
Kin Wai Ng, Francesco Antici, Nigel Tan, Befikir Bogale, Caleb Han, Florence Tama, Osamu Miyashita, Bogdan Nicolae, Michela Taufer
HPDC3
2025 ParaDyMS: Parallel Dynamic Motif Counting at Scale
abstract
Motifs in graphs (or networks) are subgraphs induced by a small set of vertices such as triangles and cliques. The frequency of motifs is used to compare and align networks across various domains, such as biology, epidemiology, and social sciences. Recent advances have made it feasible to solve the computational challenge of counting motifs in networks with over a billion edges. However, these algorithms apply only for static networks where the structure remains unchanged. In reality, networks dynamically evolve, and understanding how motifs change with the dynamic nature of these networks remains an unsolved challenge despite its potential to provide essential insights into the system. We present ParaDyMS (Parallel Dynamic Motif Counting at Scale), the first parallel algorithm for updating motif counts in fully dynamic networks using batched updates. Our algorithm updates the frequencies of motifs only in the modified parts of the network instead of recomputing them from scratch. We provide proof of the algorithm's correctness and complexity and empirically compare its execution time with another state-of-the-art static algorithm on shared memory and GPUs using realworld networks. Our results show that our algorithm is highly scalable and can significantly reduce the time to compute motifs by more than 90% in the best case and, on average, by 69%.
Nigel Tan, Jack D. Marquez, Michela Taufer, Sanjukta Bhowmick
CCGrid2
2025 On Optimizing Checkpoint Restoration for HPC Applications: Leveraging Merkle Trees and Asynchronous I/O
abstract
Efficient checkpoint restoration is critical in high-performance computing (HPC) and AI applications, where slow recovery times disrupt workflows, waste resources, and hinder reproducibility. This work introduces a Merkle tree checkpoint restoration method to accelerate failure recovery and improve explainability. Our method integrates asynchronous I/O via the Liburing library to optimize scattered reads in HPC applications. Tested on the Polaris system at Argonne National Laboratory, it exhibits lower restoration time and memory consumption than state-of-the-art checkpoint restoration methods, reaching near-full efficiency with duplicated data. Our work advances scalable and efficient checkpointing solutions for HPC, ensuring reliable and fast failure recovery for large-scale simulations.
Zackary Malkmus, Nigel Tan, Ian Lumsden, Kevin Assogba, M. Mustafa Rafique, Bogdan Nicolae, Michela Taufer
HPDC2
2024 Towards Affordable Reproducibility Using Scalable Capture and Comparison of Intermediate Multi-Run Results
abstract
Ensuring reproducibility in high-performance computing (HPC) applications is a significant challenge, particularly when nondeterministic execution can lead to untrustworthy results. Traditional methods that compare final results from multiple runs often fail because they provide sources of discrepancies only a posteriori and require substantial resources, making them impractical and unfeasible. This paper introduces an innovative method to address this issue by using scalable capture and comparing intermediate multi-run results. By capitalizing on intermediate checkpoints and hash-based techniques with user-defined error bounds, our method identifies divergences early in the execution paths. We employ Merkle trees for checkpoint data to reduce the I/O overhead associated with loading historical data. Our evaluations on the nondeterministic HACC cosmology simulation show that our method effectively captures differences above a predefined error bound and significantly reduces I/O overhead. Our solution provides a robust and scalable method for improving reproducibility, ensuring that scientific applications on HPC systems yield trustworthy and reliable results.
Nigel Tan, Kevin Assogba, Walter J. Ashworth, Befikir Bogale, Franck Cappello, M. Mustafa Rafique, Michela Taufer, Bogdan Nicolae
Middleware1
2023 Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication
abstract
Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs’ high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.
Nigel Tan, Jakob Lüttgau, Jack D. Marquez, Keita Teranishi, Nicolas M. Morales, Sanjukta Bhowmick, Franck Cappello, Michela Taufer, Bogdan Nicolae
ICPP1
2022 VPIC 2.0: Next Generation Particle-in-Cell Simulations
abstract
VPIC is a general purpose particle-in-cell simulation code for modeling plasma phenomena such as magnetic reconnection, fusion, solar weather, and laser-plasma interaction in three dimensions using large numbers of particles. VPIC's capacity in both fidelity and scale makes it particularly well-suited for plasma research on pre-exascale and exascale platforms. In this article, we demonstrate the unique challenges involved in preparing the VPIC code for operation at exascale, outlining important optimizations to make VPIC efficient on accelerators. Specifically, we show the work undertaken in adapting VPIC to exploit the portability-enabling framework Kokkos and highlight the enhancements to VPIC's modeling capabilities to achieve performance at exascale. We assess the achieved performance-portability trade-off through a suite of studies on nine different varieties of modern pre-exascale hardware. Our performance-portability study includes weak-scaling runs on three of the top ten TOP500 supercomputers, as well as a comparison of low-level system performance of hardware from four different vendors.
Robert F. Bird, Nigel Tan, Scott V. Luedtke, Stephen Lien Harrell, Michela Taufer, Brian J. Albright
IEEE Trans. Parallel Distributed Syst.2
2021 A Case Study in Scientific Reproducibility from the Event Horizon Telescope (EHT)
abstract
This poster presents the first results of an interdisciplinary project aiming to develop and share sustainable knowledge necessary to analyze, understand, and use published scientific results to advance reproducibility in multi-messenger astrophysics. Specifically, the project targets breakthrough work associated with the First M87 Event Horizon Telescope (EHT) and delivers recommendations on how the published results of the first black hole can be effectively reproduced. The project has the potential to advance new discovery in multi-messenger astrophysics by providing guidance for generalizing methods and findings from use cases.
Ross Ketron, Jacob Leonard, Brandan Roachell, Ria Patel, R. White, Silvina Caíno-Lores, Nigel Tan, Patrick R. Miles, Karan Vahi, Ewa Deelman, Duncan A. Brown, Michela Taufer
e-Science7
2021 Identifying Degree and Sources of Non-Determinism in MPI Applications Via Graph Kernels
abstract
As the scientific community prepares to deploy an increasingly complex and diverse set of applications on exascale platforms, the need to assess reproducibility of simulations and identify the root causes of reproducibility failures increases correspondingly. One of the greatest challenges facing reproducibility issues at exascale is the inherent non-determinism at the level of inter-process communication. The use of non-deterministic communication constructs is necessary to boost performance, but communication non-determinism can also hamper software correctness and result reproducibility. To address this challenge, we propose a software framework for identifying the percentage and sources of communication non-determinism. We model parallel executions as directed graphs and leverage graph kernels to characterize run-to-run variations in inter-process communication. We demonstrate the effectiveness of graph kernel similarity as a proxy for non-determinism, by showing that these kernels can quantify the type and degree of non-determinism present in communication patterns. To demonstrate our framework's ability to link and quantify runtime non-determinism to root sources, demonstrate with present for an adaptive mesh refinement application, where our framework automatically quantifies the impact of function calls on non-determinism, and a Monte Carlo application, where our framework automatically quantifies the impact of parameter configurations on non-determinism.
Dylan Chapp, Nigel Tan, Sanjukta Bhowmick, Michela Taufer
IEEE Trans. Parallel Distributed Syst.2