Mohammad Alaul Haque Monil

dblp:136/7851 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-3419-4037ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accl++ : A high-productivity programming language for performance and code portability on heterogeneous systems
abstract
This work describes the Accl++ programming language for heterogeneous computing. Accl++ is embedded in the C++ language and implemented as a C++ library. The language allows for describing device code and execution libraries and includes primitives for runtime compilation (RTC), thereby enabling code portability across disparate devices. Here, we demonstrate Accl++’s capability by coding different benchmarks from different domains. This work also describes an analysis of the overheads introduced by the Accl++ RTC support as well as the performance of the generated code for the Accl++ kernels. Accl++ improves heterogeneous code portability without incurring high levels of overhead. In two different heterogeneous systems—one composed of two 32-core AMD EPYC 7513 CPUs and two NVIDIA A100 GPUs and the other composed of two 12-core AMD EPYC 7272 CPUs and two AMD MI100 GPUs—Accl++ enables the execution of the same binary application with observed RTC overheads in the range of 3%–10% of the total kernel execution time, resulting in performance levels similar to those of native CUDA/HIP/OpenCL code.
Marc González 0001, Pedro Valero-Lara, Mohammad Alaul Haque Monil, Seyong Lee, Beau Johnston, Aaron R. Young, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter
Future Gener. Comput. Syst.3
2025 POSTER: IRISX: A Dynamic Trade-off System for Performance Portability on Multi-Accelerator Platforms
abstract
Contemporary high performance computing, cloud, and embedded systems are equipped with various combinations of heterogeneous processors. This has led to the development of portable abstractions to target these heterogeneous systems. However, these abstractions fail to deliver performance portability as they cannot easily modify key parameters such as the amount of concurrency, the representation of the computation graph, the kernel implementation, and data transfers for varying problem sizes and hardware architectures. To address this wide parameter space, this work presents IRISX, a framework that dynamically finds the performance portable solution based on application characteristics and underlying hardware. By changing the representation of the computation at the kernel level, task level, and graph level, IRISX finds the appropriate set of heterogeneous processors that provides the best performance while ensuring portability to different heterogeneous systems.
Sanil Rao, Mohammad Alaul Haque Monil, Het Mankad, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter, Franz Franchetti
PACT2
2025 IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing
abstract
In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.
Narasinga Rao Miniskar, Aaron R. Young, Mohammad Alaul Haque Monil, Kazi Asifuzzaman, Beau Johnston, Keita Teranishi, Jeffrey S. Vetter
ICPP3
2025 ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem
abstract
ChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack.
Pedro Valero-Lara, Aaron R. Young, Jeffrey S. Vetter, Zheming Jin, Swaroop Pophale, Mohammad Alaul Haque Monil, Keita Teranishi, William F. Godoy
SC6
2024 CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems
abstract
Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.
Norihisa Fujita, Beau Johnston, Narasinga Rao Miniskar, Ryohei Kobayashi 0001, Mohammad Alaul Haque Monil, Keita Teranishi, Seyong Lee, Jeffrey S. Vetter, Taisuke Boku
e-Science5
2023 Distributing Simplex-Shaped Nested for-Loops to Identify Carcinogenic Gene Combinations
abstract
Cancer is a leading cause of death in the US, and it results from a combination of two-nine genetic mutations. Identifying five-hit combinations responsible for several cancer types is computationally intractable even with the fastest super-computers in the USA. Iterating through nested loops required by the process presents a simplex-shaped workload with irregular memory access patterns. Distributing this workload efficiently across thousands of GPUs offers a challenge in dividing simplex-shaped (triangular/tetrahedral) workload into similar shapes with equal volume. Irregular memory access patterns create imbalanced compute utilization across nodes. We developed a generalized solution for distributing a simplex-shaped workload by partially coalescing the nested for-loops, minimizing the memory access overhead by efficiently utilizing limited shared memory, a dynamic scheduler, and loop tiling. For 4-hit combinations, we achieved a 90% − 100% strong scaling efficiency for up to 3594 V100 GPUs on the Summit supercomputer. Finally, we designed and implemented a distributed algorithm to identify 5-hit combinations for four different cancer types, and the identified combinations can differentiate between cancer and normal samples with 86.59−88.79% precision and 84.42 − 90.91% recall. We also demonstrated the robustness of our solution by porting the code to another leadership class computing platform Crusher, a testbed for the fastest supercomputer Frontier. On Crusher, we achieved 98% strong scaling efficiency on 50 nodes (400 AMD MI250X GCDs) and demonstrated the computational readiness of Frontier for scientific applications.
Sajal Dash, Mohammad Alaul Haque Monil, Junqi Yin, Ramu Anandakrishnan, Feiyi Wang
IPDPS2
2022 IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS Library
abstract
This paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors.
Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Pedro Valero-Lara, Frank Liu 0001, Jeffrey S. Vetter
HIPC2
2020 MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs
abstract
Integrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed.
Mohammad Alaul Haque Monil, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony
PACT1
2019 Runtime Adaptive Task Inlining on Asynchronous Multitasking Runtime Systems
abstract
As the era of high frequency, single core processors have come to a close, the new paradigm of many core processors has come to dominate. In response to these systems, asynchronous multitasking runtime systems have been developed as a promising solution to efficiently utilize these newly available hardware. Asynchronous multitasking runtime systems work by dividing a problem into a large number of fine grained tasks. However, as the number of tasks created increase, the overheads associated with task creation and management cannot be ignored. Task inlining, a method where the parent thread consumes a child thread, enables the runtime system to achieve the balance between parallelism and its overhead. As largely impacted by different processor architectures, the decision of task inlining is dynamic in nature. In this research, we present adaptive techniques for deciding, at runtime, whether a particular task should be inlined or not. We present two policies, a baseline policy that makes inlining decision based on a fixed threshold and an adaptive policy which decides the threshold dynamically at runtime. We also evaluate and justify the performance of these policies on different processor architectures. To the best of our knowledge, this is the first study of the impacts of adaptive policy at runtime for task inlining in an asynchronous multitasking runtime system on different processor architectures. From experimentation, we find that the baseline policy improves the execution time from 7.61% to 54.09%. Furthermore, the adaptive policy improves over the baseline policy by up to 74%.
Bibek Wagle, Mohammad Alaul Haque Monil, Kevin A. Huck, Allen D. Malony, Adrian Serio, Hartmut Kaiser
ICPP2
2018 Stingray-HPC: A Scalable Parallel Seismic Raytracing System
abstract
The Stingray raytracer was developed for marine seismology to compute minimum travel time from all sources in an earth model to determine the 3D geophysical structure below the ocean floor. The original sequential implementation of Stingray used Dijkstra's single-source, shortest-path (SSSP) algorithm. A data parallel version of Stingray was developed based on the Bellman-Ford-Moore iterative SSSP algorithm. Single node experiments demonstrated performance improvements from parallelization with multicore (using OpenMP) and manycore processors (using CUDA). Calculating seismic ray paths for larger earth models requires distributed, multi-node algorithms utilizing domain decomposition methods. Preliminary 2D decomposition strategies show promising scaling results. However, a general 3D decomposition methodology is needed to handle any seismic raytracing problem on any HPC computing platform. In this paper, we present Stingray-HPC, a framework for scalable seismic raytracing which can automatically decompose a 3D earth model across nodes in a distributed environment, allocate ghost cell regions for iterative updates, coordinate ghost cell communications, and test for global convergence. Stingray-HPC is implemented with MPI and either OpenMP or CUDA for node- level calculations. Our results validate Stingray-HPC's ability to handle large models (over a billion points) and to solve these models efficiently at scale up to 512 GPU nodes.
Mohammad Alaul Haque Monil, Allen D. Malony, Douglas Toomey, Kevin A. Huck
PDP1
2017 QoS-Aware Virtual Machine Consolidation in Cloud Datacenter
abstract
With the rapid growth of the cloud industry in recent years, energy consumption of warehouse-scale datacenters has become a major concern. Energy-aware Virtual Machine consolidation has proven to be one of the most effective solutions for tackling this problem. Among the sub-problems of VM consolidation, VM placement is the trickiest and can be treated as a bin packing problem which is NP-hard, hence, it is logical to apply a heuristic approach. The main challenge of VM consolidation is to achieve a balance between energy consumption and quality of service (QoS). In this research, we evaluate this problem and design a combined strategy using best fit decreasing bin packing method and multi-pass optimization in VM placement for an efficient VM consolidation. We have used CloudSim toolkit to simulate our experiments. To evaluate the performance of the proposed algorithms we used real-world workload traces from thousand VMs. Results demonstrate that our proposed methods outperform other existing methods.
Mohammad Alaul Haque Monil, Allen D. Malony
IC2E1
2015 Implementation of modified overload detection technique with VM selection strategies based on heuristics and migration control
abstract
Due to the fast growing computing need catered by Cloud infrastructure, energy consumption reduction has been an active field of research. Dynamic VM consolidation is one such research area where server load is monitored and intelligent decision is taken to optimize the usage of datacenters. In our research, we have designed an overload detection method and several VM selection methods for VM migration to achieve the trade off in VM consolidation. We have redesigned and updated an existing overload detection algorithm using heuristics and implement it in our new VM selection algorithms combining migration control and heuristics that was introduced in[9,10]. We have evaluated our algorithms through simulation on large-scale experiments driven by real-world workload traces from more than a thousand Planet lab VMs. From the comparison with existing energy aware VM consolidation methods, it is found that performance of proposed method incurs least energy consumption.
Mohammad Alaul Haque Monil, Rashedur M. Rahman
ICIS1
2013 Speed and direction based fuzzy handover system
abstract
Handover is a very common event in Cellular Mobile network; however QoS could be severely affected by the handover performance. More handover cause more signaling traffic. Therefore, it is desired that handover should be done only when it is necessary. Besides, handover decision should be precise by taking account all possible options and considering the best one. For moving Mobile Station (MS) handover takes place more frequently. Some fuzzy logic based methodologies have already been proposed in literature to provide handover decision. In this paper, a method has been proposed to calculate speed and direction of MS relative to base station as a single metric using measurement data. Accordingly, a fuzzy logic based handover algorithm is implemented to avoid ping pong effect. By taking relative speed and direction, traffic load, signal strength and distance, the fuzzy inference system determines the best candidate neighbor based on the measurement reports from MS. Simulation has been carried out in Matlab environment and a comparison of different approaches has been performed. Simulation results demonstrate that proposed algorithm provides prediction based handover decision more accurately and avoid unnecessary handover and ping pong effect.
Mohammad Alaul Haque Monil, Romasa Qasim, Rashedur M. Rahman
FUZZ-IEEE1