EDBT 2026 Demo / reviewers in the wild / expert
Ali TehraniJamsaz
dblp:314/9530 · also Ali Tehrani Jamsaz
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0001-3678-5730ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ConTraPh: Contrastive Learning for Parallelization and Performance OptimizationabstractWith the advancement of HPC platforms, the demand for high-performing applications continues to grow.One effective way to enhance program performance is through parallelization.However, fully leveraging the powerful hardware of HPC platforms poses significant challenges.Even experienced developers must carefully consider factors such as runtime, memory usage, and thread-scheduling overhead.Additionally, achieving successful parallelization often requires running applications to determine the optimal configurations.In this paper, we propose ConTraPh, a framework that integrates Contrastive Learning with Transformers and Graph Neural Networks to capture the inherent parallel characteristics of source programs through a multi-view program representation, utilizing both source code and compiler intermediate representations.This contrastive learning framework allows the model to effectively learn correct parallel configurations from positive samples while avoiding incorrect ones through negative samples.We evaluate Con-TraPh on six downstream tasks involving three different parallel programming models OpenMP, OpenCL and, Ope-nACC that include OpenMP clause prediction, performant reduction style detection, performant scheduling type detection, CPU/GPU parallelism prediction, Heterogeneous Device Mapping for OpenCL code, and OpenACC clause prediction.ConTraPh outperforms state-of-the-art models in these tasks, achieving accuracy improvements of up to 8%, 10%, 7%, 4%, 2%, and 9%, respectively.ConTraPh achieves speedups as high as 13x, 18x, 14x, and 4.4x on the reduction Quazi Ishtiaque Mahmud, Ali TehraniJamsaz, Nesreen K. Ahmed, Theodore L. Willke, Ali Jannesari |
ICS | 2 |
| 2025 | AutoParLLM: GNN-guided Context Generation for Zero-Shot Code Parallelization using LLMsabstractQuazi Ishtiaque Mahmud, Ali TehraniJamsaz, Hung D Phan, Le Chen, Mihai Capotă, Theodore L. Willke, Nesreen K. Ahmed, Ali Jannesari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Quazi Ishtiaque Mahmud, Ali TehraniJamsaz, Hung D. Phan, Mihai Capota, Theodore L. Willke, Nesreen K. Ahmed, Ali Jannesari |
NAACL (Long Papers) | 2 |
| 2024 | MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IRabstractIdentifying errors in parallel MPI programs is a challenging task. Despite the growing number of verification tools, debugging parallel programs remains a significant challenge. This paper is the first to utilize embedding and deep learning graph neural networks (GNNs) to tackle the issue of identifying bugs in MPI programs. Specifically, we have designed and developed two models that can determine, from a code’s LLVM Intermediate Representation (IR), whether the code is correct or contains a known MPI error.We tested our models using two dedicated MPI benchmark suites for verification: MBI and MPI-CorrBench. By training and validating our models on the same benchmark suite, we achieved a prediction accuracy of 92% in detecting error types. Additionally, we trained and evaluated our models on distinct benchmark suites (e.g., transitioning from MBI to MPI-CorrBench) and achieved a promising accuracy of over 80%. Finally, we investigated the interaction between different MPI errors and quantified our models generalization capabilities over new unseen errors. This involved removing errors types during training and assessing whether our models could still predict them. The detection accuracy of removed errors vary significantly between 20% to 80%, indicating connected error patterns. Jad El Karchi, Hanze Chen, Ali TehraniJamsaz, Ali Jannesari, Mihail Popov, Emmanuelle Saillard |
IPDPS | 3 |
| 2024 | CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel ProgrammingabstractAutomatic translation of programming languages has garnered renewed interest, driven by recent advancements in large language models (LLMs). Encoder-decoder transformer models, in particular, have shown promise in translating between different programming languages. However, translating between a language and its high-performance computing (HPC) extension remains underexplored due to inherent challenges like complex parallel semantics understanding. In this paper, we introduce CodeRosetta, an encoder-decoder transformer model explicitly designed for translating between programming languages and also their HPC extensions. CodeRosetta is evaluated on C++ to CUDA and Fortran to C++ translation.
It employs a customized learning-based framework with tailored pretraining and training objectives that enable it to effectively capture code semantics and parallel structural nuances, allowing for bidirectional code translation. Our results show that CodeRosetta outperforms state-of-the-art baselines in C++ to CUDA translation by 2.9 BLEU and 1.72 CodeBLUE points while improving compilation accuracy by 6.05%. Compared to general closed-source LLMs, our proposed bidirectional learning-based method improves C++ to CUDA translation by 22.08 BLEU and 14.39 CodeBLUE with 2.75% higher compilation accuracy.
Finally, CodeRosetta exhibits proficiency in Fortran to parallel C++ translation, marking it, to our knowledge, as the first encoder-decoder model for such a complex translation task, improving CodeBLEU at least by 4.63 points compared to closed-source LLMs and Open Code LLM. Ali TehraniJamsaz, Arijit Bhattacharjee, Nesreen K. Ahmed, Amir Yazdanbakhsh, Ali Jannesari |
NeurIPS | 1 |
| 2023 | Performance Optimization using Multimodal Modeling and Heterogeneous GNNabstractGrowing heterogeneity and configurability in HPC architectures has made auto-tuning applications and runtime parameters on these systems very complex. Users are presented with a multitude of options to configure parameters. In addition to application specific solutions, a common approach is to use general purpose search strategies, which often might not identify the best configurations or their time to convergence is a significant barrier. There is, thus, a need for a general purpose and efficient tuning approach that can be easily scaled and adapted to various tuning tasks. We propose a technique for tuning parallel code regions that is general enough to be adapted to multiple tasks. In this paper, we analyze IR-based programming models to make task-specific performance optimizations. To this end, we propose the Multimodal Graph Neural Network and Autoencoder (MGA) tuner, a multimodal deep learning based approach that adapts Heterogeneous Graph Neural Networks and Denoising Autoencoders for modeling IR-based code representations that serve as separate modalities. This approach is used as part of our pipeline to model a syntax, semantics, and structure-aware IR-based code representation for tuning parallel code regions/kernels. We extensively experiment on OpenMP and OpenCL code regions/kernels obtained from PolyBench, Rodinia, STREAM, DataRaceBench, AMD SDK, NPB, NVIDIA SDK, Parboil, SHOC, LULESH, XSBench, RSBench, miniFE, miniAMR, and Quicksilver benchmarks and applications. We apply our multimodal learning techniques to the tasks of (i) optimizing the number of threads, scheduling policy and chunk size in OpenMP loops and, (ii) identifying the best device for heterogeneous device mapping of OpenCL kernels. Our experiments show that this multimodal learning based approach outperforms the state-of-the-art in almost all experiments. Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo César, Anna Sikora, Ali Jannesari |
HPDC | 3 |
| 2023 | PERFOGRAPH: A Numerical Aware Program Graph Representation for Performance Optimization and Program AnalysisabstractThe remarkable growth and significant success of machine learning have expanded its applications into programming languages and program analysis. However, a key challenge in adopting the latest machine learning methods is the representation of programming languages which has a direct impact on the ability of machine learning methods to reason about programs. The absence of numerical awareness, aggregate data structure information, and improper way of presenting variables in previous representation works have limited their performances. To overcome the limitations and challenges of current program representations, we propose a novel graph-based program representation called PERFOGRAPH. PERFOGRAPH can capture numerical information and the aggregate data structure by introducing new nodes and edges. Furthermore, we propose an adapted embedding method to incorporate numerical awareness.
These enhancements make PERFOGRAPH a highly flexible and scalable representation that can effectively capture programs' intricate dependencies and semantics. Consequently, it serves as a powerful tool for various applications such as program analysis, performance optimization, and parallelism discovery. Our experimental results demonstrate that PERFOGRAPH outperforms existing representations and sets new state-of-the-art results by reducing the error rate by 7.4% (AMD dataset) and 10% (NVIDIA dataset) in the well-known Device Mapping challenge. It also sets new state-of-the-art results in various performance optimization tasks like Parallelism Discovery and Numa and Prefetchers Configuration prediction. Ali TehraniJamsaz, Quazi Ishtiaque Mahmud, Nesreen K. Ahmed, Ali Jannesari |
NeurIPS | 1 |
| 2022 | Learning Intermediate Representations using Graph Neural Networks for NUMA and Prefetchers OptimizationabstractThere is a large space of NUMA and hardware prefetcher configurations that can significantly impact the performance of an application. Previous studies have demonstrated how a model can automatically select configurations based on the dynamic properties of the code to achieve speedups. This paper demonstrates how the static Intermediate Representation (IR) of the code can guide NUMA/prefetcher optimizations without the prohibitive cost of performance profiling. We propose a method to create a comprehensive dataset that includes a diverse set of intermediate representations along with optimum configurations. We then apply a graph neural network model in order to validate this dataset. We show that our static intermediate representation based model achieves 80 % of the performance gains provided by expensive dynamic performance profiling based strategies. We further develop a hybrid model that uses both static and dynamic information. Our hybrid model achieves the same gains as the dynamic models but at a reduced cost by only profiling 30 % of the programs. Ali TehraniJamsaz, Mihail Popov, Akash Dutta, Emmanuelle Saillard, Ali Jannesari |
IPDPS | 1 |
| 2021 | Building representative and balanced datasets of OpenMP parallel regionsabstractIncorporating machine learning into automatic performance analysis and tuning tools is a promising path to tackle the increasing heterogeneity of current HPC applications. However, this introduces the need for generating balanced and representative datasets of parallel applications' executions. This work proposes a methodology for building datasets of OpenMP parallel code regions patterns. It allows for determining whether a given code region covers a unique part of the pattern input space not covered by the patterns already included in the dataset. The proposed methodology uses hardware performance counters to represent the execution of the region, which is referred to as the region signature for a given number of cores. Then, a complete representation of the region is built by joining the signatures for every different thread configuration in the system. Next, correlation analysis is performed between this representation and the representation of all the patterns already in the training set. Finally, if this correlation is below a given threshold, the region is considered to cover a unique part of the pattern input space and is subsequently added to the dataset. For validating this methodology, an example dataset, obtained from well known benchmarks, has been used to train a carefully designed neural network model to demonstrate that it is able to classify different patterns of OpenMP parallel regions. Jordi Alcaraz, Steven Sleder, Ali TehraniJamsaz, Anna Sikora, Ali Jannesari, Joan Sorribes, Eduardo César |
PDP | 3 |