Md Nahid Newaz

dblp:298/8552 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-5273-3869ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Locality Aware Process Remapping for Distributed-Memory Graph Workloads
abstract
Distributed-memory graph applications are dominated by communication and synchronization overheads. For such applications, the communication pattern comprises of variable-sized data exchanges between process neighbors in a process graph topology. Unlike process grid for rectangular problems, it is much more difficult to optimize communication for the graph topology. Custom process assignment can improve the communication performance irrespective of the data partitioning strategy. Existing automated solutions are scarce and only caters to a cartesian process topology and not the graph topology which is induced by graph-based workloads. In this paper, we propose automated network-agnostic locality-aware process assignment heuristics for distributedmemory graph workloads, based on the structure of input graphs. For four communication intensive distributed-memory graph workloads - Breadth First Search (BFS), Louvain Clustering, Triangle Counting and Single Source Shortest Path (SSSP), we demonstrate up to$30-40 \%$improvements in the overall MPI communication times through proposed process remapping methodologies via packet-level simulations using Structural Simulation Toolkit (SST) and validate the strategies empirically on HPE Slingshot network of the NERSC Perlmutter supercomputer.
Md Nahid Newaz, Nathan R. Tallent, Guangzhi Qu
IPDPS1
2024 Graph Analytics on Jellyfish topology
abstract
Because large unstructured datasets are important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the "best" network topologies for irregular communication, unstructured graph-based interconnects can be more suitable, due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.In addition to two common stencil-based mini-applications (LULESH and Sweep3D), we analyze three popular graph workloads – clustering, pattern enumeration (triangle counting), and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, considering relevant network routing schemes. Using packet-level simulations, we report average improvements of about 4–20% and 3–26% between equivalent Jellyfish vs. Dragonfly and Jellyfish vs. Fat tree topologies across diverse input graphs and applications.
Md Nahid Newaz, Joshua Suetterlein, Nathan R. Tallent, Md Atiqul Mollah, Ming Hua 0003
IPDPS1
2023 Memory Usage Prediction of HPC Workloads Using Feature Engineering and Machine Learning
abstract
In High Performance Computing (HPC) systems, numerous applications of varying scale and domain are scheduled to run concurrently, and share the available CPU and memory capacities among themselves. Applications whose run-time memory usage are not known a priori, are commonly allocated with significantly higher amounts of memory than actually needed, which leads to poor resource utilization and performance degradation of the overall system. In this paper, we disseminate our experience of performing user analysis and prediction over a large-scale resource utilization dataset to tightly estimate the memory requirements of a wide variety of applications in the Titan supercomputer system. By coupling our engineered features with random forest and XGBoost supervised machine learning techniques, our models respectively predict the correct class of memory usage in 89% and 90% of the validation data. Furthermore, more than 98% of users have 95% or better average prediction accuracy within one class tolerance range of the actual memory usage.
Md Nahid Newaz, Md Atiqul Mollah
HPC Asia1
2021 Improving adaptive routing performance on large scale Megafly topology
abstract
The Megafly topology has recently been proposed as an efficient, hierarchical way to interconnect large-scale high performance computing systems. Megafly networks may be constructed in various group sizes and configurations, but it is challenging to maintain high throughput performance on all such variants. Therefore, a robust topology-specific adaptive routing scheme is needed to utilize the topological advantages of Megafly. Currently, Progressive Adaptive Routing (PAR) is the best known routing scheme for Megafly networks, but its performance is not fully known across all scales and configurations. In this work, we show that the current PAR scheme performs sub-optimally on Megafly networks with a large number of groups. As better alternatives, we propose a new practical adaptive routing schemes, KU-GCN, that can improve the communication performance of Megafly at any scale and configuration. With the use of trace-driven simulation experiments, we show that our new Megafly routing scheme performs well across a wide variety of topology-workload combinations and outperforms PAR by up to 43.5 percent on a topology with a large number of groups.
Md Nahid Newaz, Md Atiqul Mollah, Peyman Faizian, Zhou Tong
CCGRID1
2021 Optimizing k-path selection for randomized interconnection networks
abstract
Several new interconnect designs have been proposed in the recent past for high performance computing clusters and data centers that use random connections of endpoints. Despite their superior flexibility, scalability and cost-effectiveness over conventional fat-tree based designs, practical deployment of random topologies can be prohibitive as they are prone to throughput bottlenecks if used with conventional shortest-path routing schemes. Even with multi-path routing, bottleneck issues remain on random topologies due to practical limits of routing table size. In this work, we propose a novel heuristic-driven scheme to select k best paths from all available short paths with an objective to minimize routing path bottlenecks. Our proposed scheme relies on the topological information to choose paths for each possible communication node pairs and can be paired with congestion-aware adaptive routing schemes. It, however, does not require awareness of the global traffic pattern in real time and as such, results in a stable routing control plane even under dynamic traffic conditions. We perform comparative performance analysis of our path selection with k-shortest path selection scheme on random regular graph networks and under various traffic conditions. Our experiments show that multi-path routing schemes paired with our path selection achieve up to 42 percent faster communication than the same schemes paired with k-shortest path selection, when the value of k is limited due to capacity constraints on routing table memory.
Md Nahid Newaz, Md Atiqul Mollah
HiPC1