EDBT 2026 Demo / reviewers in the wild / expert
Nicholas Malaya
dblp:134/5140
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-6259-7453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Insights from Optimizing HPL Performance on Exascale Systems: A Comparative Analysis of Panel FactorizationabstractHigh performance LINPACK (HPL) remains the primary benchmark for evaluating supercomputing performance. It includes many parts with substantial internal complexity, and its performance is affected by a large number of parameters that interact in ways that are difficult to predict on large-scale heterogeneous supercomputer systems. We present a comprehensive performance analysis of HPL on Frontier, the world’s first exascale supercomputer, which achieved HPL performance of 1.35 exaflops. Through empirical parameter tuning, detailed modeling, and comparative evaluation, we uncover critical performance insights, share lessons learned, and outline best practices for effective parameter tuning on exascale systems. We introduce and evaluate two novel PDFACT strategies: a dedicated-thread (DT) variant and a GPU-based variant (GPUPDFACT) implementation using HIP cooperative groups, demonstrating that GPU-based factorization outperforms conventional CPU-based PDFACT on Frontier’s architecture. Our findings establish key performance factors for HPL on exascale systems and offer valuable guidance for future high-performance computing and benchmarking efforts. Hao Lu 0001, Michael A. Matheson, Noel Chalmers, Aditya Kashi, Nicholas Malaya, Feiyi Wang |
SC | 5 |
| 2024 | Realizing the AMD Exascale Heterogeneous Processor Vision : Industry ProductabstractAMD had previously detailed its exascale research journey from initial targets and requirements to the development and evolution of its vision of a high-performance computing (HPC) accelerated processing unit (APU), dubbed the Exascale Heterogeneous Processor or EHP. At the conclusion of that work, the learnings were integrated into the design of the node architecture that went into the Frontier supercomputer, the world’s first exascale machine. However, while the Frontier node architecture embodied many of the attributes of the EHP concept, advanced heterogeneous integration capabilities at the time were not yet sufficiently mature to realize our vision of a fully-integrated APU for HPC and AI. In this paper, we finish the EHP’s story by digging deeper into why an APU was not the right solution at the time of our first exascale architecture, what the shortcomings were of previous EHP concepts, and how AMD further evolved the concept into the AMD Instinct™ MI300A APU. MI300A is the culmination of years of AMD developments in advanced packaging technologies, its APU hardware and software, and the next step in our highly effective chiplet strategy to not only deliver a groundbreaking design for exascale computing, but to also meet the demands of new large-language model and generative AI applications. Alan Smith 0003, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Samuel Naffziger, Mike Mantor, Nathan Kalyanasundharam, Vamsi Alla, Nicholas Malaya, Joseph L. Greathouse, Eric Chapman, Raja Swaminathan |
ISCA | 9 |
| 2023 | ADARNet: Deep Learning Predicts Adaptive Mesh RefinementabstractDeep Learning (DL) algorithms have gained popularity for super-resolution tasks - reconstructing a high-resolution (HR) output from its low-resolution (LR) counterpart. However, current DL approaches, both in computer vision and computational fluid dynamics (CFD), perform spatially uniform super-resolution. Therefore, DL for CFD approaches often over-resolve regions of the LR input that are already accurate at low numerical precision. This hardware over-utilization limits their scalability. To address this limitation, we propose ADARNet, a DL-based adaptive mesh refinement (AMR) framework. ADARNet takes a LR image as input and outputs its non-uniform HR counterpart, predicting HR only in areas that require higher numerical accuracy. As a result, ADARNet predicts the target 1024 × 1024 solution 7 − 28.5 × faster than state-of-the-art DL methods and reduces the memory usage by 4.4 − 7.65 × while maintaining the same level of accuracy. Moreover, unlike traditional AMR solvers that refine the mesh iteratively, ADARNet is a one-shot method that accelerates it by 2.6 − 4.5 ×. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
ICPP | 3 |
| 2023 | A Research Retrospective on AMD's Exascale Computing JourneyabstractThe pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing. Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov |
ISCA | 44 |
| 2023 | Experiences readying applications for ExascaleabstractThe advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems. Nicholas Malaya, O. E. Bronson Messer, Joseph Glenski, Antigoni Georgiadou, Justin Lietz, Kalyana C. Gottiparthi, Marcus S. Day, Jackie Chen, Jon S. Rood, Lucas Esclapez, James B. White III, Gustav R. Jansen, Nicholas Curtis, Stephen Nichols, Jakub Kurzak, Noel Chalmers, Chip Freitag, Paul T. Bauman, Alessandro Fanfarillo, Reuben D. Budiardja, Thomas Papatheodore, Nicholas Frontiere, Damon McDougall, Matthew R. Norman, Sarat Sreepathi, Philip C. Roth, Dmytro Bykov, Noah Wolfe, Paul Mullowney, Markus Eisenbach 0002, Marc T. Henry de Frahan, Wayne Joubert |
SC | 1 |
| 2021 | SURFNet: Super-Resolution of Turbulent Flows with Transfer Learning using Small DatasetsabstractDeep Learning (DL) algorithms are emerging as a key alternative to computationally expensive CFD simulations. However, state-of-the-art DL approaches require large and high-resolution training data to learn accurate models. The size and availability of such datasets are a major limitation for the development of next-generation data-driven surrogate models for turbulent flows. This paper introduces SURFNet, a transfer learning-based super-resolution flow network. SURFNet primarily trains the DL model on low-resolution datasets and transfer learns the model on a handful of high-resolution flow problems-accelerating the traditional numerical solver independent of the input size. We propose two approaches to transfer learning for the task of super-resolution, namely one-shot and incremental learning. Both approaches entail transfer learning on only one geometry to account for fine-grid flow fields requiring 15× less training data on high-resolution inputs compared to the tiny resolution ($64\times 256$) of the coarse model significantly, reducing the time for both data collection and training. We empirically evaluate SURFNet's performance by solving the Navier-Stokes equations in the turbulent regime on input resolutions up to 256× larger than the coarse model. On four test geometries and eight flow configurations unseen during training, we observe a consistent 2–2.1× speedup over the OpenFOAM physics solver independent of the test geometry and the resolution size (up to$2048 \times 2048$), demonstrating both resolution-invariance and generalization capabilities. Moreover, compared to the baseline model (aka oracle) that collects large training data at$256 \times 256$and$512 \times 512$grid resolutions, SURFNet achieves the same performance gain while reducing the combined data collection and training time by 3.6× and 10.2×, respectively. Our approach addresses the challenge of reconstructing high-resolution solutions from coarse grid models trained using low-resolution inputs (i.e., super-resolution) without loss of accuracy and requiring limited computational resources. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
PACT | 3 |
| 2020 | CFDNet: a deep learning-based accelerator for fluid simulationsabstractCFD is widely used in physical system design and optimization, where it is used to predict engineering quantities of interest, such as the lift on a plane wing or the drag on a motor vehicle. However, many systems of interest are prohibitively expensive for design optimization, due to the expense of evaluating CFD simulations. Octavi Obiols-Sales, Abhinav Vishnu, Nicholas Malaya, Aparna Chandramowlishwaran |
ICS | 3 |
| 2017 | Accelerating Matrix Processing with GPUsabstractMatrix operations are common and expensive computations in a variety of applications. They occur frequently in high-performance computing, graphics, graph processing, and machine learning applications. This paper discusses how to map a variety of important matrix computations, including sparse matrix-vector multiplication (SpMV), sparse triangle solve (SpTS), graph processing, and dense matrix-matrix multiplication, to GPUs. Since many emerging systems will use heterogeneous architectures (e.g. CPUs and GPUs) to attain the desired performance targets under strict power constraints, this paper discusses implications and future research for matrix processing with heterogeneous designs. Conclusions common to the matrix operations discussed in this paper are: (1) Future algorithms should be written to ensure that the essential computations fit into local memory, which may require direct programmer management. (2) Algorithms are needed that expose high levels of parallelism. (3) While the scale of computation is often sufficient to support algorithms with superior asymptotic order, additional considerations, such as memory capacity and bandwidth, must also be carefully managed. (4) Libraries should be used to provide portable performance. Nicholas Malaya, Shuai Che, Joseph L. Greathouse, René van Oostrum, Michael J. Schulte |
ARITH | 1 |
| 2013 | Petascale direct numerical simulation of turbulent channel flow on up to 786K coresabstractWe present results of performance optimization for direct numerical simulation (DNS) of wall bounded turbulent flow (channel flow). DNS is a technique in which the fluid flow equations are solved without subgrid modeling. Of particular interest are high Reynolds number (Re) turbulent flows over walls, because of their importance in technological applications. Simulating high Re turbulence is a challenging computational problem, due to the high spatial and temporal resolution requirements. Myoungkyu Lee, Nicholas Malaya, Robert D. Moser |
SC | 2 |