Khalid Hasanov

dblp:131/6718 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
1since 2021 · last 2025
0009-0003-3200-7768ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Programming languages and type systems
webassembly
0.912025
Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming Languages · HPDC 2025
Performance modeling and evaluation
benchmarking
0.912025
Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming Languages · HPDC 2025

Methods — techniques the papers use, named apart from their topics

cross-language comparison · 2.6benchmarking · 2.6
YearPublicationVenuePosition
2025 Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming Languages
abstract
WebAssembly (WASM) has emerged as a promising compilation target for languages traditionally not executed on the web and enable the cross-platform deployment of high-performance applications. While its general use cases have been well studied, the performance implications of executing Machine Learning (ML) workloads via WASM across different programming languages and runtime environments remain relatively unexplored. This paper presents a systematic evaluation of two representative ML models, K-Means and Logistic Regression, implemented in Python, Rust, and C++ and compiled to WASM. These models are executed in two distinct environments: a web browser and the WebAssembly System Interface (WASI), and their execution time and accuracy is compared against each programming language across both environments. This study aims to provide insight into the trade-offs and practical considerations involved in deploying ML workloads using WebAssembly across different language ecosystems and runtime configurations.
Sallar Khan, Tania Malik, Khalid Hasanov
HPDC3
2018 A taxonomy of task-based parallel programming technologies for high-performance computing
abstract
Task-based programming models for shared memory—such as Cilk Plus and OpenMP 3—are well established and documented. However, with the increase in parallel, many-core, and heterogeneous systems, a number of research-driven projects have developed more diversified task-based support, employing various programming and runtime features. Unfortunately, despite the fact that dozens of different task-based systems exist today and are actively used for parallel and high-performance computing (HPC), no comprehensive overview or classification of task-based technologies for HPC exists. In this paper, we provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms. We demonstrate the usefulness of our taxonomy by classifying state-of-the-art task-based environments in use today.
Peter Thoman, Kiril Dichev, Thomas Heller, Roman Iakymchuk, Xavier Aguilar, Khalid Hasanov, Philipp Gschwandtner, Pierre Lemarinier, Stefano Markidis, Herbert Jordan, Thomas Fahringer, Kostas Katrinis, Erwin Laure, Dimitrios S. Nikolopoulos
J. Supercomput.6
2017 Hierarchical redesign of classic MPI reduction algorithms
Khalid Hasanov, Alexey L. Lastovetsky
J. Supercomput.1
2016 Architecting Malleable MPI Applications for Priority-driven Adaptive Scheduling
abstract
Future supercomputers will need to support both traditional HPC applications and Big Data/High Performance Analysis applications seamlessly in a common environment. This motivates traditional job scheduling systems to support malleable jobs along with allocations that can dynamically change in size, in order to adapt the amount of resources to the actual current need of the different applications. It also calls for future innovative HPC applications to adapt to this environment, and provide some level of malleability for releasing underutilized resources to other tasks. In this paper, we present and compare two different methodologies to support such malleable MPI applications: 1)using checkpoint/restart and the SCR library, and 2) using dynamic data redistribution and the ULFM API and runtime. We examine their effects on application execution times as well as their impact on resource management.
Pierre Lemarinier, Khalid Hasanov, Srikumar Venugopal, Kostas Katrinis
EuroMPI2
2015 Asymmetric communication models for resource-constrained hierarchical ethernet networks
abstract
Summary Communication time prediction is critical for parallel application performance tuning, especially for the rapidly growing field of data‐intensive applications. However, making such predictions accurately is non‐trivial when contention exists on different components in hierarchical networks. In this article, we derive an ‘asymmetric network property’ on transmission control protocol (TCP) layer for concurrent bidirectional communications in a commercial off‐the‐shelf (COTS) cluster and develop a communication model as the first effort to characterize the communication times on hierarchical Ethernet networks with contentions on both network interface card and backbone cable levels. We develop a micro‐benchmark for a set of simultaneous point‐to‐point message‐passing interface (MPI) operations on a parametrized network topology and use it to validate our model extensively and show that the model can be used to predict the communication times for simultaneous MPI operations (both point‐to‐point and collective communications) on resource‐constrained networks effectively. We show that if the asymmetric network property is excluded from the model, the communication time predictions will be significantly less accurate than those made by using the asymmetric network property. In addition, we validate the model on a cluster of Grid5000 infrastructure, which is a more loosely coupled platform. As such, we advocate the potential to integrate this model in performance analysis for data‐intensive parallel applications. Our observation of the performance degradation caused by the asymmetric network property suggests that some part of the software stack below TCP layer in COTS clusters needs targeted tuning, which has not yet attracted any attention in literature. Copyright © 2014 John Wiley & Sons, Ltd.
Jun Zhu 0003, Alexey L. Lastovetsky, Shoukat Ali, Rolf Riesen, Khalid Hasanov
Concurr. Comput. Pract. Exp.5
2015 Hierarchical approach to optimization of parallel matrix multiplication on large-scale platforms
Khalid Hasanov, Jean-Noël Quintin, Alexey L. Lastovetsky
J. Supercomput.1
2013 Hierarchical Parallel Matrix Multiplication on Large-Scale Distributed Memory Platforms
abstract
Matrix multiplication is a very important computation kernel both in its own right as a building block of many scientific applications and as a popular representative for other scientific applications. Cannon's algorithm which dates back to 1969 was the first efficient algorithm for parallel matrix multiplication providing theoretically optimal communication cost. However this algorithm requires a square number of processors. In the mid-1990s, the SUMMA algorithm was introduced. SUMMA overcomes the shortcomings of Cannon's algorithm as it can be used on a nonsquare number of processors as well. Since then the number of processors in HPC platforms has increased by two orders of magnitude making the contribution of communication in the overall execution time more significant. Therefore, the state of the art parallel matrix multiplication algorithms should be revisited to reduce the communication cost further. This paper introduces a new parallel matrix multiplication algorithm, Hierarchical SUMMA (HSUMMA), which is a redesign of SUMMA. Our algorithm reduces the communication cost of SUMMA by introducing a two-level virtual hierarchy into the two-dimensional arrangement of processors. Experiments on an IBM Blue Gene/P demonstrate the reduction of communication cost up to 2.08 times on 2048 cores and up to 5.89 times on 16384 cores.
Jean-Noël Quintin, Khalid Hasanov, Alexey L. Lastovetsky
ICPP2