EDBT 2026 Demo / reviewers in the wild / expert
Tan Nguyen 0001
dblp:74/9775-1
· DBLP profile ↗
13ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0003-3748-403XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hierarchical Methodology for Hardware Design Comparison in HPC WorkloadsabstractAs Moore's law slows down, developers face difficult choices between low-level HDLs (Verilog, VHDL) offering fine-grained control and higher-level tools (HLS and Chisel) promising improved productivity. While high-level tools accelerate development, performance gaps persist compared to expert HDL implementations. Prior studies emphasize end-to-end performance, offering limited insight into why tools excel or where performance diverges in the design hierarchy. We introduce a hierarchical framework for comparing hardware generation tools by decomposing HPC kernels (FFT, GEMM, QR factorization) into reusable primitives (MAC arrays, butterflies, permutations, reduction trees). Across Verilog, Chisel, and Vivado HLS, we built an automated tool flow and synthesized ~1, 500 variants on AMD Alveo U250, measuring resource utilization and frequency. We derived theoretical bounds for validation. Verilog achieves the highest frequency and lowest resource usage; Chisel performs comparably (5--15% gap), while HLS shows a 20--40% gap. All tools operate within bounds for well-structured designs. Crucially, performance divergence arises during primitive assembly, indicating that high-level tools require better composition optimization. This reproducible framework provides actionable insights and is extensible to other tools, domains, and FPGA architectures. Doru-Thom Popovici, Mario Vega, Angelos Ioannou, Fabien Chaix, Dania Susanne Mosuli, Blair Reasoner, Tan Nguyen 0001, Xiaokun Yang, John Shalf |
FPGA | 7 |
| 2026 | CAC: An asynchronous non-blocking consistency model with bounded staleness for distributed machine learningabstractRelaxed consistency models have been reported to significantly improve the performance of training machine learning (ML) models compared to strong consistency models. However, the existing relaxed consistency models for distributed ML either force fast workers to wait for stragglers (e.g., stale-synchronous models) or have no upper bound on data staleness (e.g., asynchronous models), thus negatively affecting the quality and performance of training ML models. We propose a new asynchronous non-blocking consistency model with bounded staleness for distributed ML that overcomes the drawbacks of blocking and unbounded staleness in previous relaxed consistency models. The new model, named Chained Asynchronous Consistency (CAC), guarantees an upper bound on data staleness in asynchronous computation without forcing fast workers to wait for slow workers. We theoretically prove that the Stochastic Gradient Descent (SGD) algorithm under CAC converges and the upper bound on the convergence expectation is independent of the number of workers. Based on the CAC model, we develop a new staleness-aware cacheable distributed object (CAC-object) for distributed ML where shared parameters are distributed among workers in a peer-to-peer manner. This approach avoids intermediate centralized storage, such as parameter servers, while maintaining simple one-sided communication (e.g., get and put ). The CAC-object allows remote parameters to be cached and reused locally following a consistency model (e.g., CAC, stale-synchronous or asynchronous model). To demonstrate the applicability of the CAC-object in asynchronous distributed ML, we introduce a new asynchronous distributed matrix completion algorithm (CAC-MF) using the CAC-object. We develop the CAC-object and CAC-MF using UPC++, a Partitioned Global Address Space (PGAS) library for high-performance computing (HPC), and evaluate them in different execution scenarios (e.g., with and without stragglers and crash failures, different minibatch sizes) on HPC clusters. Our experimental results show that the CAC model scales with increasing workers, tolerates stragglers and crash failures, and achieves better convergence than stale-synchronous and asynchronous models. Particularly, in the case of halted stragglers, CAC’s Root Mean Square Error (RMSE) is up to 25 times less than the baseline models’ RMSE for the Netflix dataset, and 170 times less for the MovieLens dataset. Phuong Hoai Ha, Xing Cai, Tan Nguyen 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | Uniconn: A Uniform High-Level Communication Library for Portable Multi-GPU ProgrammingabstractModern HPC and AI systems increasingly rely on multi-GPU clusters, where communication libraries such as MPI, NCCL/RCCL, and NVSHMEM enable data movement across GPUs. While these libraries are widely used in frameworks and solver packages, their distinct APIs, synchronization models, and integration mechanisms introduce programming complexity and limit portability. Performance also varies across workloads and system architectures, making it difficult to achieve consistent efficiency. These issues present a significant obstacle to writing portable, high-performance code for large-scale GPU systems. We present Uniconn, a unified, portable high-level C++ communication library that supports both point-to-point and collective operations across GPU clusters. Uniconn enables seamless switching between backends and APIs (host or device) with minimal or no changes to application code. We describe its design and core constructs, and evaluate its performance using network benchmarks, a Jacobi solver, and a Conjugate Gradient solver. Across three supercomputers, we compare Uniconn's overhead against CUDA/ROCm-aware MPI, NCCL/RCCL, and NVSHMEM on up to 64 GPUs. In most cases, Uniconn incurs negligible overhead, typically under 1 % for the Jacobi solver and under 2% for the Conjugate Gradient solver. Dogan Sagbili, Sinan Ekmekçibasi, Khaled Z. Ibrahim, Tan Nguyen 0001, Didem Unat |
CLUSTER | 4 |
| 2024 | Devastator: A Scalable Parallel Discrete Event Simulation Framework for Modern C++abstractParallel discrete event simulation is a fundamental simulation technology that is essential to the parallelization of event-based models including hardware and transportation systems. Parallelization is often difficult due to dynamic data-dependencies and limited computational work for hiding runtime overheads. We present Devastator, a scalable parallel discrete event simulation framework for modern C++. Devastator provides a productive API that leverages C++’s type system to eliminate boilerplate code. Devastator has been designed to specifically optimize performance on distributed many-core architectures with deepening memory hierarchies. Devastator relies on the GASNet-Ex communication runtime as well as lock-free message queues to achieve highly competitive performance. We perform strong and weak scaling studies on NERSC’s Perlmutter up to 32,698 cores and demonstrate an up to 5x speed-up over the ROSS simulator for workloads with substantial locality. John Bachan, Jianlan Ye, Tan Nguyen 0001, Mahesh Natarajan, Maximilian H. Bremer, Cy P. Chan |
SIGSIM-PADS | 4 |
| 2023 | Benefits of Optimistic Parallel Discrete Event Simulation for Network-on-Chip SimulationabstractThe end of Moore's law has placed a two-fold demand on hardware simulation. Firstly, efficient co-design requires fast simulation of hardware systems in order to vet proposed designs. Secondly, modern simulator platforms need to become increasingly concurrent as well. To address these challenges, we develop an optimistic time warp-based parallelization for the Structural Simulation Toolkit (SST). Our optimistic PDES engine hides synchronization costs by speculatively executing tasks, leading to better compute resource utilization on modern multicore architectures. Given the significant engineering effort to make custom existing hardware models reversible, we also develop a new SST component, called escher. The escher workflow instruments arbitrary SST applications and replays event traces through the different SST parallelizations. This is a useful tool for hardware engineers who want to understand the benefits of optimistic parallelization before undertaking the significant software engineering effort to make their hardware models reversible. We demonstrate this workflow by generating traces for a tiled mesh-NOC architecture and show a 2.1x to 3.7x speed-up using our optimistic SST parallelization versus the current conservative SST parallelization. Maximilian H. Bremer, Nirmalendu Bikash Patra, Tan Nguyen 0001, Dilip P. Vasudevan, Cy P. Chan |
DS-RT | 3 |
| 2022 | FPGA-based HPC accelerators: An evaluation on performance and energy efficiencyabstractAbstract Hardware specialization is a promising direction for the future of digital computing. Reconfigurable technologies enable hardware specialization with modest non‐recurring engineering cost, but their performance and energy efficiency compared to state‐of‐the‐art processor architectures remain an open question. In this article, we use FPGAs to evaluate the benefits of building specialized hardware for numerical kernels found in scientific applications. In order to properly evaluate performance, we not only compare Intel Arria 10 and Xilinx U280 performance against Intel Xeon, Intel Xeon Phi, and NVIDIA V100 GPUs, but we also extend the Empirical Roofline Toolkit (ERT) to FPGAs in order to assess our results in terms of the Roofline model. We show design optimization and tuning techniques for peak FPGA performance at reasonable hardware usage and power consumption. As FPGA peak performance is known to be far less than that of a GPU, we also benchmark the energy efficiency of each platform for the scientific kernels comparing against microbenchmark and technological limits. Results show that while FPGAs struggle to compete in absolute terms with GPUs on memory‐ and compute‐intensive kernels, they require far less power and can deliver nearly the same energy efficiency. Tan Nguyen 0001, Colin MacLean, Marco Siracusa, Douglas Doerfler, Nicholas J. Wright, Samuel Williams 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Phase asynchronous AMR execution for productive and performant astrophysical flows
Muhammed Nufail Farooqi, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf, Didem Unat |
SC | 2 |
| 2017 | Nonintrusive AMR Asynchrony for Communication Optimization
Muhammed Nufail Farooqi, Didem Unat, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf |
Euro-Par | 3 |
| 2017 | Automatic translation of MPI source into a latency-tolerant, data-driven form
Tan Nguyen 0001, Pietro Cicotti, Eric J. Bylaska, Daniel J. Quinlan, Scott B. Baden |
J. Parallel Distributed Comput. | 1 |
| 2016 | Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinementabstractHardware architecture is increasingly complex, urging the development of asynchronous runtime systems with advance resource and locality management supports. However, these supports may come at the cost of complicating the user interface while programming remains one of the major constraints to wide adoption of asynchronous runtimes in practice. In this paper, we propose a solution that leverages application metadata to enable challenging optimizations as well as to facilitate the task of transforming legacy code to an asynchronous representation. We develop Perilla, a task graph-based runtime system that requires only modest programming effort. Perilla utilizes metadata of an AMR software framework to enable various optimizations at the communication layer without complicating its API. Experimental results with different applications on up to 24K processor cores show that Perilla can realize up to 1.44x speedup over the synchronous code variant. The metadata enabled optimizations account for 25% to 100% of the performance improvement. Tan Nguyen 0001, Didem Unat, Weiqun Zhang, Ann S. Almgren, Muhammed Nufail Farooqi, John Shalf |
SC | 1 |
| 2015 | LU Factorization: Towards Hiding Communication Overheads with a Lookahead-Free AlgorithmabstractLookahead is a well-known technique for masking communication in matrix factorization, but at the cost of complicating application software. We present a new approach, based on automated code-restructuring, that realizes the benefits of lookahead while avoiding the complications. We apply our technique to HPL, the Linpack benchmark used to assess the performance of supercomputers. Starting with the simpler non-lookahead version of the application, we are able to meet the performance of lookahead on the Stampede mainframe. Tan Nguyen 0001, Scott B. Baden |
CLUSTER | 1 |
| 2013 | A software-based dynamic-warp scheduling approach for load-balancing the Viola-Jones face detection algorithm on GPUs
Tan Nguyen 0001, Daniel Hefenbrock, Jason Oberg, Ryan Kastner, Scott B. Baden |
J. Parallel Distributed Comput. | 1 |
| 2012 | Bamboo: translating MPI applications to a latency-tolerant, data-driven formabstractWe present Bamboo, a custom source-to-source translator that transforms MPI C source into a data-driven form that automatically overlaps communication with available computation. Running on up to 98304 processors of NERSC's Hopper system, we observe that Bamboo's overlap capability speeds up MPI implementations of a 3D Jacobi iterative solver and Cannon's matrix multiplication. Bamboo's generated code meets or exceeds the performance of hand optimized MPI, which includes split-phase coding, the method classically employed to hide communication. We achieved our results with only modest amounts of programmer annotation and no intrusive reprogramming of the original application source. Tan Nguyen 0001, Pietro Cicotti, Eric J. Bylaska, Dan Quinlan, Scott B. Baden |
SC | 1 |