VLDB 2026 Research / reviewers in the wild / expert
Tania Malik
dblp:155/5487
· DBLP profile ↗
4ranked-venue papers
3as first author
2since 2021 · last 2026
0000-0002-4461-7120ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 67% Performance modeling and evaluation · 26% GPUs and heterogeneous computing · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
data-parallel programming |
1.0 | 1 | 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication Links · IEEE Trans. Parallel Distributed Syst. 2026 |
Parallel and multicore computing › data distribution
domain partitioning |
1.0 | 1 | 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication Links · IEEE Trans. Parallel Distributed Syst. 2026 |
Parallel and multicore computing
load balancing |
1.0 | 1 | 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication Links · IEEE Trans. Parallel Distributed Syst. 2026 |
Programming languages and type systems
webassembly |
0.9 | 1 | 2025 | Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming Languages · HPDC 2025 |
Performance modeling and evaluation
benchmarking |
0.9 | 1 | 2025 | Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming Languages · HPDC 2025 |
Performance modeling and evaluation › cost modeling
communication cost modeling |
0.3 | 1 | 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication Links · IEEE Trans. Parallel Distributed Syst. 2026 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous processors |
0.3 | 1 | 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication Links · IEEE Trans. Parallel Distributed Syst. 2026 |
Methods — techniques the papers use, named apart from their topics
cross-language comparison · 2.6benchmarking · 2.6simulation · 1.0analytical modeling · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimal Partitioning of Square Computational Domains for Parallel Computing on Hybrid Servers With Heterogeneous Processors and Heterogeneous Communication LinksabstractIn this work, we formulate and solve the mathematical problem of optimal partitioning of a real-valued square computational domain across three heterogeneous processors connected by three heterogeneous communication links. The objective is to minimize the communication time of the data-parallel application processing this domain, assuming that its computation time is minimized by balancing the load of the processors. The state-of-the-art methods ignore the heterogeneity of the communication links and try to minimize the total amount of communicated data instead of the communication time. As a result, optimal partitions found by these methods may not minimize the communication time on mainstream heterogeneous hybrid servers with heterogeneous data links. We introduce a bandwidth-aware communication cost function, representing the communication time of parallel processing for each given partition. Using this function, we identify and prove the optimality of four partitioning shapes and 24 candidate partitions, which can potentially minimize the communication cost. We derive analytical formulas for the communication cost of each candidate partition and use them to find the optimal one for each given combination of bandwidths of communication links and ratio of processor speeds. The state-of-the-art bandwidth-oblivious methods only identify 3 optimal partitioning shapes and 3 candidate partitions. Through extensive simulations, we compare optimal partitioning shapes found by the bandwidth-aware and bandwidth-oblivious methods for the full range of processor speed ratios and realistic bandwidths of communication links. We also experimented on a real hybrid heterogeneous server comprising a multi-core CPU and two accelerators. These experiments demonstrate the superior prediction accuracy of the bandwidth-aware communication model over its bandwidth-oblivious counterpart, resulting in truly optimal partitioning solutions in all conducted experiments. Tania Malik, Alexey L. Lastovetsky |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Performance Evaluation of Machine Learning Applications Using WebAssembly Across Different Programming LanguagesabstractWebAssembly (WASM) has emerged as a promising compilation target for languages traditionally not executed on the web and enable the cross-platform deployment of high-performance applications. While its general use cases have been well studied, the performance implications of executing Machine Learning (ML) workloads via WASM across different programming languages and runtime environments remain relatively unexplored. This paper presents a systematic evaluation of two representative ML models, K-Means and Logistic Regression, implemented in Python, Rust, and C++ and compiled to WASM. These models are executed in two distinct environments: a web browser and the WebAssembly System Interface (WASI), and their execution time and accuracy is compared against each programming language across both environments. This study aims to provide insight into the trade-offs and practical considerations involved in deploying ML workloads using WebAssembly across different language ecosystems and runtime configurations. Sallar Khan, Tania Malik, Khalid Hasanov |
HPDC | 2 |
| 2020 | Optimal Matrix Partitioning for Data Parallel Computing on Hybrid Heterogeneous PlatformsabstractIn this paper, we study the problem of partitioning a matrix over a small number of interconnected heterogeneous processors. This problem is crucial for data parallel dense linear algebra and other applications with similar communication patterns on modern hybrid servers, integrating several heterogeneous compute devices such as CPUs, GPUs and other accelerators. The objective is to balance the load of the heterogeneous devices while minimising the communication cost. While the problem has been solved for the case of two processors, it is still open for three and more processors. The state-of-the-art solution for the case of three processors uses a communication cost function, which does not accurately account for the total amount of data moved between processors and therefore leaves the question of its global optimality open. In this work, we propose a cost function, which accurately represents the total amount of data moved between processors. Then, we formulate and solve the problem of optimal partitioning of a square computational domain, using this accurate communication cost function. Finally, we propose and implement an original experimental methodology for accurate measurement of the communication time of parallel applications on hybrid heterogeneous servers, integrating multi-core CPUs and various accelerators. We apply this methodology to experimental validation of our mathematical result. Tania Malik, Alexey L. Lastovetsky |
ISPDC | 1 |
| 2016 | Network-aware optimization of communications for parallel matrix multiplication on hierarchical HPC platformsabstractSummary Communications on hierarchical heterogeneous high‐performance computing platforms can be optimized based on topology and performance information. For MPI, as a major programming tool for such platforms, a number of topology‐aware and performance‐aware implementations of collective operations have been proposed for optimal scheduling of messages. This approach improves performance of application and does not require to modify application source code. However, it is applicable to collective operations only and does not affect the parts of the application that are based on point‐to‐point exchanges. In this paper, we address the problem of efficient execution of data‐parallel applications on interconnected clusters and present optimizations that improve data partition by taking into account the entire communication flow of the application. This approach is also non‐intrusive to the source code but application specific. For illustration, we use parallel matrix multiplication, where the matrices are partitioned into irregular two‐dimensional rectangles assigned to different processors and arranged in columns, and the processors communicate over this partition vertically and horizontally. By rearranging the rectangles, we can minimize communications between different levels of the network hierarchy. Finding the optimal arrangement is NP‐complete; therefore, we propose two heuristic approaches based on evaluation of the communication flow on the given network topology. We demonstrate the correctness and efficiency of the proposed approaches by experimental results on multicore nodes and interconnected heterogeneous clusters. Copyright © 2015 John Wiley & Sons, Ltd. Tania Malik, Vladimir Rychkov, Alexey L. Lastovetsky |
Concurr. Comput. Pract. Exp. | 1 |