Md Atiqul Mollah

dblp:170/2092 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-2789-0846ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Graph Analytics on Jellyfish topology
abstract
Because large unstructured datasets are important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the "best" network topologies for irregular communication, unstructured graph-based interconnects can be more suitable, due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.In addition to two common stencil-based mini-applications (LULESH and Sweep3D), we analyze three popular graph workloads – clustering, pattern enumeration (triangle counting), and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, considering relevant network routing schemes. Using packet-level simulations, we report average improvements of about 4–20% and 3–26% between equivalent Jellyfish vs. Dragonfly and Jellyfish vs. Fat tree topologies across diverse input graphs and applications.
Md Nahid Newaz, Joshua Suetterlein, Nathan R. Tallent, Md Atiqul Mollah, Ming Hua 0003
IPDPS5
2023 Memory Usage Prediction of HPC Workloads Using Feature Engineering and Machine Learning
abstract
In High Performance Computing (HPC) systems, numerous applications of varying scale and domain are scheduled to run concurrently, and share the available CPU and memory capacities among themselves. Applications whose run-time memory usage are not known a priori, are commonly allocated with significantly higher amounts of memory than actually needed, which leads to poor resource utilization and performance degradation of the overall system. In this paper, we disseminate our experience of performing user analysis and prediction over a large-scale resource utilization dataset to tightly estimate the memory requirements of a wide variety of applications in the Titan supercomputer system. By coupling our engineered features with random forest and XGBoost supervised machine learning techniques, our models respectively predict the correct class of memory usage in 89% and 90% of the validation data. Furthermore, more than 98% of users have 95% or better average prediction accuracy within one class tolerance range of the actual memory usage.
Md Nahid Newaz, Md Atiqul Mollah
HPC Asia2
2021 Improving adaptive routing performance on large scale Megafly topology
abstract
The Megafly topology has recently been proposed as an efficient, hierarchical way to interconnect large-scale high performance computing systems. Megafly networks may be constructed in various group sizes and configurations, but it is challenging to maintain high throughput performance on all such variants. Therefore, a robust topology-specific adaptive routing scheme is needed to utilize the topological advantages of Megafly. Currently, Progressive Adaptive Routing (PAR) is the best known routing scheme for Megafly networks, but its performance is not fully known across all scales and configurations. In this work, we show that the current PAR scheme performs sub-optimally on Megafly networks with a large number of groups. As better alternatives, we propose a new practical adaptive routing schemes, KU-GCN, that can improve the communication performance of Megafly at any scale and configuration. With the use of trace-driven simulation experiments, we show that our new Megafly routing scheme performs well across a wide variety of topology-workload combinations and outperforms PAR by up to 43.5 percent on a topology with a large number of groups.
Md Nahid Newaz, Md Atiqul Mollah, Peyman Faizian, Zhou Tong
CCGRID2
2021 Optimizing k-path selection for randomized interconnection networks
abstract
Several new interconnect designs have been proposed in the recent past for high performance computing clusters and data centers that use random connections of endpoints. Despite their superior flexibility, scalability and cost-effectiveness over conventional fat-tree based designs, practical deployment of random topologies can be prohibitive as they are prone to throughput bottlenecks if used with conventional shortest-path routing schemes. Even with multi-path routing, bottleneck issues remain on random topologies due to practical limits of routing table size. In this work, we propose a novel heuristic-driven scheme to select k best paths from all available short paths with an objective to minimize routing path bottlenecks. Our proposed scheme relies on the topological information to choose paths for each possible communication node pairs and can be paired with congestion-aware adaptive routing schemes. It, however, does not require awareness of the global traffic pattern in real time and as such, results in a stable routing control plane even under dynamic traffic conditions. We perform comparative performance analysis of our path selection with k-shortest path selection scheme on random regular graph networks and under various traffic conditions. Our experiments show that multi-path routing schemes paired with our path selection achieve up to 42 percent faster communication than the same schemes paired with k-shortest path selection, when the value of k is limited due to capacity constraints on routing table memory.
Md Nahid Newaz, Md Atiqul Mollah
HiPC2
2018 A Comparative Study of Topology Design Approaches for HPC Interconnects
abstract
The recent interconnect topology designs for High Performance Computing (HPC) systems have followed two directions, one characterized by low diameter and the other by high path diversity. The low diameter design focuses on building large networks with small diameters, guaranteeing one short path between each pair of nodes. Examples include Slim Fly and Dragonfly. The high path diversity design takes into account not only other topological metrics such as diameter but also path diversity between pairs of nodes. Examples include fat-tree, Random Regular Graph (RRG) and Generalized De Bruin Graph (GDBG). Topologies designed from these two approaches have distinct features and require very different routing schemes to exploit the network capacity. In this work, we study the performance-related topological features of representative topologies of the two design approaches, including Slim Fly, Dragonfly, RRG, and GDBG, and compare HPC application performance on these topologies with a set of routing schemes. The study uncovers new knowledge about the topologies designed by these two approaches. Findings of the study include (1) the load balance routing technique designed for low diameter topologies, known as the Universal Globally Adaptive Load-balanced routing (UGAL), can be effectively adapted for the high path diversity topologies, and (2) high path diversity topologies in general achieve higher performance than low diameter topologies for networks built by a similar number of the same type of switches.
Md Atiqul Mollah, Peyman Faizian, Md. Shafayat Rahman, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
CCGrid1
2018 Load-Balanced Slim Fly Networks
abstract
The Slim Fly topology has recently been proposed for the future generation supercomputers. It has small diameter and relies on the Universal Globally Adaptive Load-balanced (UGAL) routing, which adapts the routes between minimal (MIN) routing and Valiant Load-Balancing (VLB) routing to exploit the network capacity. In this work, we show that the current Slim Fly is not load-balanced for both MIN routing and VLB routing, in that certain links in the network have a significantly higher probability to carry traffic than others. As such, hot spots are more likely to form on such links. We propose two approaches to address this problem and to make Slim Fly load-balanced: (1) modifying the topology by selectively increasing the bandwidth of the potential hot-spot links so that the original routing becomes load-balanced, and (2) modifying the routing scheme by using a weighted VLB routing to distribute the traffic in a more load balanced fashion than the original VLB routing on the original Slim Fly. The results of our performance analysis and simulation demonstrate that both approaches result in more effective Slim Fly than its current form.
Md. Shafayat Rahman, Md Atiqul Mollah, Peyman Faizian, Xin Yuan 0001
ICPP2
2018 Random Regular Graph and Generalized De Bruijn Graph with k-Shortest Path Routing
abstract
The Random regular graph (RRG) has recently been proposed as an interconnect topology for future large scale data centers and HPC clusters. An RRG is a special case of directed regular graph (DRG) where each link is unidirectional and all nodes have the same number of incoming and outgoing links. In this work, we establish bounds for DRGs on diameter, average k-shortest path length, and a load balancing property with k-shortest path routing, and use these bounds to evaluate RRGs. The results indicate that an RRG with k-shortest path routing is not ideal in terms of diameter and load balancing. We further consider the Generalized De Bruijn Graph (GDBG), a deterministic DRG, and prove that for most network configurations, a GDBG is near optimal in terms of diameter, average k-shortest path length, and load balancing with a k-shortest path routing scheme. Finally, we use modeling and simulation to exploit the strengths and weaknesses of RRGs for different traffic conditions by comparing RRGs with GDBGs.
Peyman Faizian, Md Atiqul Mollah, Xin Yuan 0001, Zaid Salamah A. Alzaid, Scott Pakin, Michael Lang 0003
IEEE Trans. Parallel Distributed Syst.2
2018 Rapid Calculation of Max-Min Fair Rates for Multi-Commodity Flows in Fat-Tree Networks
abstract
Max-min fairness is often used in the performance modeling of interconnection networks. Existing methods to compute max-min fair rates for multi-commodity flows have high complexity and are computationally infeasible for large networks. In this work, we show that by considering topological features, this problem can be solved efficiently for the fat-tree topology that is widely used in data centers and high performance compute clusters. Several efficient new algorithms are developed for this problem, including a parallel algorithm that can take advantage of multi-core and shared-memory architectures. Using these algorithms, we demonstrate that it is possible to find the max-min fair rate allocation for multi-commodity flows in fat-tree networks that support tens of thousands of nodes. We evaluate the run-time performance of the proposed algorithms and show improvement in orders of magnitude over the previously best known method. We further demonstrate a new application of max-min fair rate allocation that is only computationally feasible using our new algorithms.
Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
IEEE Trans. Parallel Distributed Syst.1
2017 A comparative study of SDN and adaptive routing on dragonfly networks
abstract
The OpenFlow-style Software Defined Networking (SDN) technology has shown promising performance in data centers and campus networks; and the HPC community is significantly interested in adopting the SDN technology. However, while OpenFlow-style SDN allows dynamic per-flow resource management using a global network view, it does not support adaptive routing, which is widely used in HPC systems. This gives rise to the question whether SDN can achieve the performance that HPC systems expect with adaptive routing. In this work, we investigate possible methods to apply the SDN technology on the current generation HPC interconnects with the Dragonfly topology, and compare the performance of SDN with that of adaptive routing. Our results indicate that adaptive routing results in higher performance than SDN when both have similar resource allocation for a given traffic condition. However, SDN can use the global network view to compete with adaptive routing by allocating network resources more effectively.
Peyman Faizian, Md Atiqul Mollah, Zhou Tong, Xin Yuan 0001, Michael Lang 0003
SC2
2016 Random Regular Graph and Generalized De Bruijn Graph with k-Shortest Path Routing
abstract
Random regular graph (RRG) has recently been proposed as an interconnect topology for future large scale data centers and HPC clusters. While various studies have been performed, this topology is still not well understood. RRG is a special case of directed regular graph (DRG) where each link is unidirectional and all nodes have the same number of incoming and outgoing links. In this work, we establish bounds for DRG on diameter, average k-shortest path length, and a load balancing property with k-shortest path routing, and use these bounds to evaluate RRG. The results indicate that RRG with k-shortest path routing is not ideal in terms of diameter and load balancing. We further consider the Generalized De Bruijn Graph (GDBG), a deterministic DRG, and prove that for most network configurations, GDBG is near optimal in terms of diameter, average k-shortest path length, and load balancing with a k-shortest path routing scheme. Finally, we explore the strengths and weaknesses of RRG for different traffic conditions by comparing RRG with GDBG.
Peyman Faizian, Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
IPDPS2
2015 Fast Calculation of Max-Min Fair Rates for Multi-commodity Flows in Fat-Tree Networks
abstract
Max-min fairness is often used in the performance modeling of interconnection networks. Existing methods to compute max-min fair rates for multi-commodity flows have high complexity and are computationally infeasible for large networks. In this work, we show that by considering topological features, this problem can be solved efficiently for the fat-tree topology that is widely used in data centers and high performance computing clusters. Using two new algorithms that we developed, we demonstrate it is possible to find the max-min fair rate allocation for multi-commodity flows in fat-tree networks that support tens of thousands of nodes. We evaluate the run-time performance of the proposed algorithms and demonstrate an application.
Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
CLUSTER1