Ali Jannesari

dblp:74/1277 · also Ali Jannesari Ladani · DBLP profile ↗
← Back
75ranked-venue papers
6as first author
44since 2021 · last 2026
0000-0001-8672-5317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 54 · 5 first-author · 29 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
abstract
Le Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nuo Xu 0013, Winson Chen, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao
ACL (1)9
2026 Dual-Distilled Heterogeneous Federated Learning with Adaptive Margins for Trainable Global Prototypes
Fatema Siddika, Md. Anwar Hossen, Anuj Sharma 0001, Juan Pablo Muñoz, Ali Jannesari
CCGrid6
2026 SuperSFL: Resource-Heterogeneous Federated Split Learning with Weight-Sharing Supernet
Abdullah Al Asif, Sixing Yu, Juan Pablo Muñoz, Arya Mazaheri, Ali Jannesari
Euro-Par (2)5
2026 Resource-Aware Online Tuning for Heterogeneous Federated Learning
abstract
In federated learning, differences in client data, compute capability, memory, and network conditions make training with a single global configuration difficult. We propose a resource-aware, agent-based online tuning framework that adapts hyperparameters during training using client resource profiles and recent learning dynamics. Unlike wrapper-based approaches, the method performs tuning while training, avoiding repeated end-to-end retraining and adding minimal runtime overhead. Across representative federated benchmarks, it consistently improves accuracy, achieving 5–10% gains over strong baselines, while being designed to support efficient training under heterogeneous client conditions.
Abdullah Al Asif, Juan Pablo Muñoz, Ali Jannesari
HPDC4
2026 A Lock-Free Work-Stealing Algorithm for Bulk Operations
abstract
Work-stealing is a widely used technique for balancing irregular parallel workloads, and most modern runtime systems adopt lock-free work-stealing deques to reduce contention and improve scalability. However, existing algorithms are designed for general-purpose parallel runtimes and often incur overheads that are unnecessary in specialized settings. In this paper, we present a new lock-free work-stealing queue tailored for a master-worker framework used in the parallelization of a mixed-integer programming optimization solver based on decision diagrams. Our design supports native bulk operations, grows without bounds, and assumes at most one owner and one concurrent stealer, thereby eliminating the need for heavy synchronization.
Raja Sai Nandhan Yadav Kataru, Danial Davarnia, Ali Jannesari
HPDC3
2026 Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
abstract
Large-scale Graph Neural Networks (GNNs) are typically trained by sampling a vertex’s neighbors to a fixed distance. Because large input graphs are distributed, training requires frequent irregular communication that stalls forward progress. Moreover, fetched data changes with graph, graph distribution, sample and batch parameters, and caching policies. Consequently, any static prefetching method will miss crucial opportunities to adapt to different dynamic conditions.
Aishwarya Sarkar, Nathan R. Tallent, Aman Chadha, Tanya G. Roosta, Ali Jannesari
ICS6
2026 Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications
abstract
With the growing prevalence of heterogeneous computing, CPUs are increasingly being paired with accelerators to achieve new levels of performance and energy efficiency. However, data movement between devices remains a significant bottleneck, complicating application development. Existing performance tools require considerable programmer intervention to diagnose and locate data transfer inefficiencies. To address this, we propose dynamic analysis techniques to detect and profile inefficient data transfer and allocation patterns in heterogeneous applications. We implemented these techniques into OMPDataPerf, which provides detailed traces of problematic data mappings, source code attribution, and assessments of optimization potential in heterogeneous OpenMP applications. OMPDataPerf uses the OpenMP Tools Interface (OMPT) and incurs only a 5 % geometric‑mean runtime overhead.
Luke J. Marzen, Junhyung Shim, Ali Jannesari
PPoPP3
2025 Federated Multimodal Learning with Dual Adapters and Selective Pruning for Communication and Computational Efficiency
abstract
Federated Learning (FL) enables collaborative learning across distributed clients while preserving data privacy. However, FL faces significant challenges when dealing with heterogeneous data distributions, which can lead to suboptimal global models that fail to generalize across diverse clients. In this work, we propose a novel framework designed to tackle these challenges by introducing a dual-adapter approach. The method utilizes a larger local adapter for client-specific personalization and a smaller global adapter to facilitate efficient knowledge sharing across clients. Additionally, we incorporate a pruning mechanism to reduce communication overhead by selectively removing less impactful parameters from the local adapter. Through extensive experiments on a range of vision and language tasks, our method demonstrates superior performance compared to existing approaches. It achieves higher test accuracy, lower performance variance among clients, and improved worst-case performance, all while significantly reducing communication and computation costs. Overall, the proposed method addresses the critical trade-off between model personalization and generalization, offering a scalable solution for real-world FL applications.
Duy Phuong Nguyen, Juan Pablo Muñoz, Tanya G. Roosta, Ali Jannesari
CCGrid4
2025 HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction
abstract
Accurate flood forecasting remains a critical challenge for water-resource management, as it demands simultaneous modeling of local, time-varying runoff drivers (e.g., rainfall-induced peaks, base- flow trends) and complex spatial interactions across a river network. Traditional data-driven approaches, such as convolutional networks and sequence-based models, ignore topological information about the region. Graph Neural Networks (GNNs), in contrast, propagate information exactly along the river network, making them ideal for learning hydrological routing. However, state-of-the-art GNN-based flood prediction models still collapse pixels to coarse catchment polygons because the cost of training explodes with graph size and higher resolution. Furthermore, most existing methods treat spatial and temporal dependencies separately, either applying GNNs solely on spatial graphs or transformers purely on temporal sequences, thus failing to simultaneously capture spatiotemporal interactions critical for accurate flood prediction. To address these limitations, we introduce a heterogenous basin graph to represent every land and river pixel as a node connected by both physical hydrological flow directions as well as inter-catchment relationships. We also propose HydroGAT, a novel spatiotemporal network that adaptively learns both local temporal importance as well as most influential upstream locations. Evaluated in two Midwestern US basins and across five baseline architectures, our model achieves higher NSE (up to 0.97), improved KGE (up to 0.96), and low bias (PBIAS within ± 5%) in hourly discharge prediction, while offering interpretable attention maps that reveal sparse, structured intercatchment influences. To support high-resolution basin-scale training, we develop a distributed data-parallel pipeline that scales efficiently up to 64 NVIDIA A100 GPUs on NERSC Perlmutter supercomputer, demonstrating up to 15× speedup across machines. Our code is available at https://github.com/swapp-lab/HydroGAT.
Aishwarya Sarkar, Autrin Hakimi, Xiaoqiong Chen, Hai Huang 0015, Chaoqun Lu, Ibrahim Demir, Ali Jannesari
SIGSPATIAL/GIS7
2025 Weight-Sharing NAS with Architecture-Agnostic Intermediate Representation
abstract
Weight-sharing supernet has been widely adopted in Neural Architecture Search (NAS) as a promising strategy to obtain smaller and more efficient high-performance models. However, constructing supernets requires domain expertise to design architecture-specific rules (e.g., rules for CNNs and Transformers) for generating subnets, and training a supernet demands joint optimization over a vast sample space of subnets, which is computationally expensive. This paper presents OSF (Optimized Supernet Formation), an automated and architecture-agnostic approach that transforms predefined/pretrained models into weight-sharing supernets. Specifically, we propose representing neural architectures using a high-level computational graph intermediate representation (IR) that enables both the conversion of different types of models into supernets and the extraction of executable subnets via graph traversal. To improve supernet training efficiency, we introduce a sampling strategy that prioritizes the most promising subnet candidates during training, and propose a fork-join parallel training approach with gradient accumulation that resolves write-after-write dependencies, enabling concurrent training of multiple subnet architectures with shared weights. Our empirical evaluations demonstrate that OSF successfully builds supernets from various architectures (CNNs, Transformers, SSMs, and MLPs) while achieving superior performance across language and vision benchmarks. Notably, for Vision Transformers (ViT), OSF reduces FLOPs by 49% while maintaining the accuracy, resulting in a 155% increase in throughput and 35% latency reduction. Code Open-sourced at: https://github.com/yusx-swapp/OSF
Sixing Yu, Arya Mazaheri, Ali Jannesari
HPDC3
2025 HHOTuner: Efficient Performance Tuning with Harris Hawks Optimization
abstract
As we enter the Post-Exascale Computing era, parallel programs are becoming more ubiquitous. Increased parallelization efforts and larger, more complex systems have led to numerous ways of tweaking performance of such applications. Existing tuners often produce very good results, but are often task specific, or have high overheads. Our aim is to enable and adapt optimization techniques seen in nature for auto-tuning HPC search spaces. We believe that such an approach can outperform the auto-tuning abilities of state-of-the-art tuners while significantly reducing associated overheads. To this end, we propose the HHOTuner (Harris Hawks Optimization Tuner): a nature-inspired meta-heuristic, swarm-based optimization technique for auto-tuning real-world HPC search spaces. Named after the Harris hawks, the HHOTuner employs and adapts the real-world predatory behavior of Harris hawks as algorithms for tuning real-world HPC applications and workloads. The proposed auto-tuner is general purpose and is designed to work well with user-defined search spaces. We have evaluated HHOTuner on several workloads such as HPC applications, graph neural network epoch training, and tensor program generation by an auto-tuning compiler. HHOTuner improves, sometimes significantly, the performance of these workloads. Its performance is better than state-of-the-art auto-tuners, sometimes being an order of magnitude ahead for a few cases. HHOTuner’s tuning overhead is much lower (up to 7.2 × and 32.9 ×) than the other auto-tuners evaluated in this paper, while maintaining or improving the quality of auto-tuning results.
Akash Dutta, Ali Jannesari
ICPP2
2025 ConTraPh: Contrastive Learning for Parallelization and Performance Optimization
abstract
With the advancement of HPC platforms, the demand for high-performing applications continues to grow.One effective way to enhance program performance is through parallelization.However, fully leveraging the powerful hardware of HPC platforms poses significant challenges.Even experienced developers must carefully consider factors such as runtime, memory usage, and thread-scheduling overhead.Additionally, achieving successful parallelization often requires running applications to determine the optimal configurations.In this paper, we propose ConTraPh, a framework that integrates Contrastive Learning with Transformers and Graph Neural Networks to capture the inherent parallel characteristics of source programs through a multi-view program representation, utilizing both source code and compiler intermediate representations.This contrastive learning framework allows the model to effectively learn correct parallel configurations from positive samples while avoiding incorrect ones through negative samples.We evaluate Con-TraPh on six downstream tasks involving three different parallel programming models OpenMP, OpenCL and, Ope-nACC that include OpenMP clause prediction, performant reduction style detection, performant scheduling type detection, CPU/GPU parallelism prediction, Heterogeneous Device Mapping for OpenCL code, and OpenACC clause prediction.ConTraPh outperforms state-of-the-art models in these tasks, achieving accuracy improvements of up to 8%, 10%, 7%, 4%, 2%, and 9%, respectively.ConTraPh achieves speedups as high as 13x, 18x, 14x, and 4.4x on the reduction
Quazi Ishtiaque Mahmud, Ali TehraniJamsaz, Nesreen K. Ahmed, Theodore L. Willke, Ali Jannesari
ICS5
2025 Adaptive Federated Distillation with Dual-LoRA for Personalized Representation Learning
abstract
In federated learning for multimodal models, balancing global generalization and local personalization remains challenging due to heterogeneous client data distributions. To address this, we propose a novel federated learning framework employing dual Low-Rank Adaptation (LoRA) adapters within a frozen Contrastive Language-Image Pre-training (CLIP) backbone. Each client maintains both global and local adapters, dynamically orchestrated via a lightweight gating network that adaptively fuses the adapters based on input-specific features. Unlike traditional parameter averaging, our server aggregates client knowledge through federated distillation on a compact reference dataset, effectively mitigating parameter conflicts across clients. Experiments demonstrate that our adaptive fusion strategy significantly improves personalized representation quality, outperforming standard LoRA-based federated approaches, especially in scenarios with diverse local data distributions.
Duy Phuong Nguyen, Chianing Johnny Wang, Ali Jannesari
SEC3
2025 PCEBench: A Multi-Dimensional Benchmark for Evaluating Large Language Models in Parallel Code Generation
abstract
The increasing complexity of software systems and advancements in hardware architectures have intensified the demand for efficient parallel code generation. While parallel programming offers significant performance benefits, it requires extensive expertise and effort due to the intricacies of synchronization, data management, and optimizations. To address these challenges, recent studies have explored the application of machine learning (ML) techniques in parallel code generation, aiming to reduce manual efforts and enhance performance outcomes. Large language models (LLMs) have recently revolutionized natural language processing (NLP) and demonstrated remarkable capabilities in code generation. However, evaluating their ability to generate high-performance parallel code presents unique challenges. Unlike sequential code evaluation, the evaluation of LLM-generated parallel code requires consideration of not only correctness but also efficiency and scalability in utilizing parallel resources. Concretely, existing benchmarks for LLM-generated parallel code evaluation are limited in size and scope compared to their sequential counterparts. To address this evaluation gap, we introduce PCEBench, a novel benchmark designed to assess LLMs' capabilities in generating parallel code. PCEBench focuses on multi-tasking and multidimensional performance evaluation, leveraging an LLM-based approach to generate verified prompts for parallel code generation. The benchmark incorporates scripts compatible with compilers and data race checkers, enabling comprehensive testing across critical dimensions such as compilability, executability, code self-correctness, functional correctness, data race detection, and speedup over serial implementations. By examining these multiple dimensions, PCEBench not only facilitates a thorough evaluation of LLMs in parallel code generation but also provides valuable insights for developers to enhance model performance in this challenging task. This comprehensive approach contributes to advancing the field of automated parallel programming and supports the development of more efficient and scalable software systems.
Nesreen K. Ahmed, Mihai Capota, Theodore L. Willke, Niranjan Hasabnis, Ali Jannesari
IPDPS6
2025 AutoParLLM: GNN-guided Context Generation for Zero-Shot Code Parallelization using LLMs
abstract
Quazi Ishtiaque Mahmud, Ali TehraniJamsaz, Hung D Phan, Le Chen, Mihai Capotă, Theodore L. Willke, Nesreen K. Ahmed, Ali Jannesari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Quazi Ishtiaque Mahmud, Ali TehraniJamsaz, Hung D. Phan, Mihai Capota, Theodore L. Willke, Nesreen K. Ahmed, Ali Jannesari
NAACL (Long Papers)8
2024 Improving Federated Learning Through Low-Entropy Client Sampling Based on Learned High-Level Features
abstract
Data heterogeneity impacts the performance of Federated Learning (FL) by introducing training noise. Although representative client sampling can help mitigate the issue, it remains challenging to implement without compromising data privacy. This work introduces a new method to address the problem by proposing an affordable blind (privacy preserving) clustering mechanism for conducting stratified client sampling. Inspired by the ‘dialect quiz’, we propose a ‘response test’ to cluster clients whose models have learned similar high-level features. This approach facilitates representative client sampling without the need for direct access to client data. We demonstrate empirically that our method yields client samples with low relative entropy with respect to the global data distribution, indicating increased representativeness. Convergence experiments reveal that applying our method significantly improves the convergence and accuracy of the global model compared to strong baselines like SCAFFOLD and FL-CIR. Additionally, the reduced number of training rounds required to achieve target accuracy leads to decreased communication overhead and computational expense, making our approach promising for practical FL implementations.
Waqwoya Abebe, Pablo Munoz, Ali Jannesari
CLOUD3
2024 MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
abstract
One of the primary areas of interest in High Performance Computing is the improvement of performance of parallel workloads. Nowadays, compilable source code-based optimization tasks that employ deep learning often exploit LLVM Intermediate Representations (IRs) for extracting features from source code. Most such works target specific tasks, or are designed with a pre-defined set of heuristics. So far, pre-trained models are rare in this domain, but the possibilities have been widely discussed. Especially approaches mimicking large-language models (LLMs) have been proposed. But these have prohibitively large training costs. In this paper, we propose MIREncoder, a Multi-modal IR-based Auto-Encoder that can be pre-trained to generate a learned embedding space to be used for downstream tasks by machine learning-based approaches. A multi-modal approach enables us to better extract features from compilable programs. It allows us to better model code syntax, semantics and structure. For code-based performance optimizations, these features are very important while making optimization decisions. A pre-trained model/embedding implicitly enables the usage of transfer learning, and helps move away from task-specific trained models. Additionally, a pre-trained model used for downstream performance optimization should itself have reduced overhead, and be easily usable. These considerations have led us to propose a modeling approach that i) understands code semantics and structure, ii) enables use of transfer learning, and iii) is small and simple enough to be easily re-purposed or reused even with low resource availability. Our evaluations will show that our proposed approach can outperform the state of the art while reducing overhead.
Akash Dutta, Ali Jannesari
PACT2
2024 MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
abstract
Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learning on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized continuous prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15-40% improvement in end-to-end training performance on the National Energy Research Scientific Computing Center's (NERSC) Perlmutter supercomputer for various OGB datasets.
Aishwarya Sarkar, Nathan R. Tallent, Ali Jannesari
CLUSTER4
2024 Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models
abstract
Foundation Models (FMs), such as LLaMA, BERT, GPT, ViT, and CLIP, have demonstrated remarkable success in a wide range of applications, driven by their ability to leverage vast amounts of data for pre-training. However, optimizing FMs often requires access to sensitive data, raising privacy concerns and limiting their applicability in many domains. In this paper, we propose the Federated Foundation Models (FFMs) paradigm, which combines the benefits of FMs and Federated Learning (FL) to enable privacy-preserving and collaborative learning across multiple end-users. We discuss the potential benefits and challenges of integrating FL into the lifespan of FMs, covering pre-training, fine-tuning, and application. We further outline potential future research avenues in FFM, including FFM pre-training, FFM fine-tuning, and federated prompt tuning, which allow the development of more personalized and context-aware models while ensuring data privacy. Moreover, we explore the possibility of continual/lifelong learning in FFMs, as increased computational power at the edge may unlock the potential for optimizing FMs using newly generated private data close to the data source. The proposed FFM concepts offer a flexible and scalable framework for training large language models in a privacy-preserving manner, setting the stage for subsequent advancements in both FM training and federated learning.
Sixing Yu, Juan Pablo Muñoz, Ali Jannesari
LREC/COLING3
2024 Leveraging Statistical Machine Translation for Code Search
abstract
Machine Translation (MT) has numerous applications in Software Engineering (SE). Recently, it has been employed not only for programming language translation but also as an oracle for deriving information for various research problems in SE. In this application branch, MT’s impact has been assessed through metrics measuring the accuracy of these problems rather than traditional translation evaluation metrics. For code search, a recent work, ASTTrans, introduced an MT-based model for extracting relevant non-terminal nodes from the Abstract Syntax Tree (AST) of an implementation based on natural language descriptions. While ASTTrans demonstrated the effectiveness of MT in enhancing code search on small datasets with low embedding dimensions, it struggled to improve the accuracy of code search on the standard benchmark CodeSearchNet.
Hung Phan, Ali Jannesari
EASE2
2024 OMPGPT: A Generative Pre-trained Transformer Model for OpenMP
Arijit Bhattacharjee, Nesreen K. Ahmed, Niranjan Hasabnis, Gal Oren 0001, Vy A. Vo, Ali Jannesari
Euro-Par (1)7
2024 Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
abstract
Abstract Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Akash Dutta, Ali Jannesari, Jordi Alcaraz
Euro-Par (1)6
2024 Resource-Aware Heterogeneous Federated Learning with Specialized Local Models
Sixing Yu, Juan Pablo Muñoz, Ali Jannesari
Euro-Par (1)3
2024 MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
abstract
Identifying errors in parallel MPI programs is a challenging task. Despite the growing number of verification tools, debugging parallel programs remains a significant challenge. This paper is the first to utilize embedding and deep learning graph neural networks (GNNs) to tackle the issue of identifying bugs in MPI programs. Specifically, we have designed and developed two models that can determine, from a code’s LLVM Intermediate Representation (IR), whether the code is correct or contains a known MPI error.We tested our models using two dedicated MPI benchmark suites for verification: MBI and MPI-CorrBench. By training and validating our models on the same benchmark suite, we achieved a prediction accuracy of 92% in detecting error types. Additionally, we trained and evaluated our models on distinct benchmark suites (e.g., transitioning from MBI to MPI-CorrBench) and achieved a promising accuracy of over 80%. Finally, we investigated the interaction between different MPI errors and quantified our models generalization capabilities over new unseen errors. This involved removing errors types during training and assessing whether our models could still predict them. The detection accuracy of removed errors vary significantly between 20% to 80%, indicating connected error patterns.
Jad El Karchi, Hanze Chen, Ali TehraniJamsaz, Ali Jannesari, Mihail Popov, Emmanuelle Saillard
IPDPS4
2024 CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
abstract
Automatic translation of programming languages has garnered renewed interest, driven by recent advancements in large language models (LLMs). Encoder-decoder transformer models, in particular, have shown promise in translating between different programming languages. However, translating between a language and its high-performance computing (HPC) extension remains underexplored due to inherent challenges like complex parallel semantics understanding. In this paper, we introduce CodeRosetta, an encoder-decoder transformer model explicitly designed for translating between programming languages and also their HPC extensions. CodeRosetta is evaluated on C++ to CUDA and Fortran to C++ translation. It employs a customized learning-based framework with tailored pretraining and training objectives that enable it to effectively capture code semantics and parallel structural nuances, allowing for bidirectional code translation. Our results show that CodeRosetta outperforms state-of-the-art baselines in C++ to CUDA translation by 2.9 BLEU and 1.72 CodeBLUE points while improving compilation accuracy by 6.05%. Compared to general closed-source LLMs, our proposed bidirectional learning-based method improves C++ to CUDA translation by 22.08 BLEU and 14.39 CodeBLUE with 2.75% higher compilation accuracy. Finally, CodeRosetta exhibits proficiency in Fortran to parallel C++ translation, marking it, to our knowledge, as the first encoder-decoder model for such a complex translation task, improving CodeBLEU at least by 4.63 points compared to closed-source LLMs and Open Code LLM.
Ali TehraniJamsaz, Arijit Bhattacharjee, Nesreen K. Ahmed, Amir Yazdanbakhsh, Ali Jannesari
NeurIPS6
2024 Work-In-Progress: Energy and Thermal-Aware Scheduling based on HMARL for OpenMP DAG Workloads
abstract
With technological advancement, the leakage power of multicore platforms which correlates exponentially with chip temperature becomes more than dynamic power. Energy-aware hardware solutions use dynamic voltage and frequency scaling (DVFS) to avoid overheating in aggressive performance-boosting scenarios, and software solutions assign high-utilization tasks to different core configurations to conserve processor package power. Since the current energy-aware heuristics are not based on core-by-core frequency monitoring, they do not address overheating when some cores are more active than others. Also, assigning tasks to cores without comprehensive profiling does not address irregular task execution. In this article, we investigate potential of applying an online reinforcement learning method to efficiently allocate tasks based on their profiling data, including makespan, to appropriate core combinations according to their temperature status, aiming to reduce energy consumption. We propose using a hierarchical multi-agent reinforcement learning (HMARL) approach, where one agent selects cores and frequency scaling as a function of profiler results, while another agent selects core combinations as a function of temperature sensor results.
Mohammad Pivezhandi, Abusayeed Saifullah, Ali Jannesari
RTSS4
2024 PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
abstract
Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques reduce bottlenecks associated with memory bandwidth, but also increase end-to-end latency per inference run, requiring high speculation acceptance rates to improve performance. Combined with a variable rate of acceptance across tasks, speculative inference techniques can result in reduced performance. Additionally, pipeline-parallel designs require many user requests to maintain maximum utilization. As a remedy, we propose PipeInfer, a pipelined speculative acceleration technique to reduce inter-token latency and improve system utilization for single-request scenarios while also improving tolerance to low speculation acceptance rates and low-bandwidth interconnects. PipeInfer exhibits up to a $2.15 \times$ improvement in generation speed over standard speculative inference. PipeInfer achieves its improvement through Continuous Asynchronous Speculation and Early Inference Cancellation, the former improving latency and generation speed by running single-token inference simultaneously with several speculative runs, while the latter improves speed and latency by skipping the computation of invalidated runs, even in the middle of inference.
Branden Butler, Sixing Yu, Arya Mazaheri, Ali Jannesari
SC4
2024 Static Generation of Efficient OpenMP Offload Data Mappings
abstract
Increasing heterogeneity in HPC architectures and compiler advancements have led to OpenMP being frequently used to enable computations on heterogeneous devices. However, the efficient movement of data on heterogeneous computing platforms is crucial for achieving high utilization. Programmers must explicitly map data between the host and connected accelerator devices to achieve efficient data movement. Ensuring efficient data transfer requires programmers to reason about complex data flow. This can be a laborious and error-prone process since the programmer must keep a mental model of data validity and lifetime spanning multiple data environments. We present a static analysis tool, OMPDart (OpenMP Data Reduction Tool), for OpenMP programs that models data dependencies between host and device regions and applies source code transformations to achieve efficient data transfer. Our evaluations on nine HPC benchmarks demonstrate that OMPDart is capable of generating effective data mapping constructs that substantially reduce data transfer between host and device.
Luke J. Marzen, Akash Dutta, Ali Jannesari
SC3
2024 Fast data-dependence profiling through prior static analysis
abstract
Data-dependence profiling is a program-analysis technique for detecting parallelism opportunities in sequential programs. It captures data dependences that actually occur during program execution, filtering parallelism-preventing dependences that purely static methods assume only because they lack critical runtime information, such as the values of pointers and array indices. Profiling, however, suffers from high runtime overhead. In our earlier work, we accelerated data-dependence profiling by excluding polyhedral loops that can be handled statically using certain compilers and eliminating scalar variables that create statically-identifiable data dependences. In this paper, we combine the two methods and integrate them into DiscoPoP, a data-dependence profiler and parallelism discovery tool. Additionally, we detect reduction patterns statically and unify the three static analyses with the DiscoPoP framework to significantly diminish the profiling overhead and for a wider range of programs. We have evaluated our unified approaches with 49 benchmarks from three benchmark suites and two computer simulation applications. The evaluation results show that our approach reports fewer false positive and negative data dependences than the original data-dependence profiler and reduces the profiling time by at least 43%, with a median reduction of 76% across all programs. Also, we identify 40% of reduction cases statically and eliminate the associated profiling overhead for these cases.
Mohammad Norouzi Arab, Nicolas Morew, Qamar Ilias, Lukas Rothenberger, Ali Jannesari, Felix Wolf 0001
Parallel Comput.5
2023 Optimizing Decentralized Learning with Local Heterogeneity using Topology Morphing and Clustering
abstract
Recently, local peer topology has been shown to influence the overall convergence of decentralized learning (DL) graphs in the presence of data heterogeneity. In this paper, we demonstrate the advantages of constructing a proxy-based locally heterogeneous DL topology to enhance convergence and maintain data privacy. In particular, we propose a novel peer clumping strategy to efficiently cluster peers before arranging them in a final training graph. By showing how locally heterogeneous graphs outperform locally homogeneous graphs of similar size and from the same global data distribution, we present a strong case for topological pre-processing. Moreover, we demonstrate the scalability of our approach by showing how the proposed topological pre-processing overhead remains small in large graphs while the performance gains get even more pronounced. Furthermore, we show the robustness of our approach in the presence of network partitions.
Waqwoya Abebe, Ali Jannesari
CCGrid2
2023 Heterogeneous Federated Learning using Dynamic Model Pruning and Adaptive Gradient
abstract
Federated Learning (FL) has emerged as a new paradigm for training machine learning models distributively without sacrificing data security and privacy. Learning models on edge devices such as mobile phones is one of the most common use cases for FL. However, Non-identical independent distributed (non-IID) data in edge devices easily leads to training failures. Especially, over-parameterized machine learning models can easily be over-fitted on such data, hence, resulting in inefficient federated learning and poor model performance. To overcome the over-fitting issue, we proposed an adaptive dynamic pruning approach for FL, which can dynamically slim the model by dropping out unimportant parameters, hence, preventing over-fittings. Since the machine learning model's parameters react differently for different training samples, adaptive dynamic pruning will evaluate the salience of the model's parameter according to the input training sample, and only retain the salient parameter's gradients when doing back-propagation. We performed comprehensive experiments to evaluate our approach. The results show that our approach by removing the redundant parameters in neural networks can significantly reduce the over-fitting issue and greatly improves the training efficiency. In particular, when training the ResNet-32 on CIFAR-10, our approach reduces the communication cost by 57%. We further demonstrate the inference acceleration capability of the proposed algorithm. Our approach reduces up to 50% FLOPs inference of DNNs on edge devices while maintaining the model's quality.
Sixing Yu, Ali Anwar 0001, Ali Jannesari
CCGrid4
2023 Evaluating and Optimizing the Effectiveness of Neural Machine Translation in Supporting Code Retrieval Models: A Study on the CAT Benchmark
abstract
Neural Machine Translation (NMT) is widely applied in software engineering tasks. The effectiveness of NMT for code retrieval relies on the ability to learn from the sequence of tokens in the source language to the sequence of tokens in the target language. While NMT performs well in pseudocode-to-code translation[17], it might have challenges in learning to translate from natural language query to source code in newly curated real-world code documentation/ implementation datasets. In this work, we analyze the performance of NMT in natural language-to-code translation in the newly curated CAT benchmark[31] that includes the optimized versions of three Java datasets TLCodeSum, CodeSearchNet, Funcom, and a Python dataset PCSD. Our evaluation shows that NMT has low accuracy, measured by CrystalBLEU[10] and Meteor[9] metrics in this task. To alleviate the duty of NMT in learning complex representation of source code, we propose ASTTrans Representation, a tailored representation of an Abstract Syntax Tree (AST) using a subset of non-terminal nodes. We show that the classical approach NMT performs significantly better in learning ASTTrans Representation over code tokens with up to 36% improvement on Meteor score. Moreover, we leverage ASTTrans Representation to conduct combined code search processes from the state-of-the-art code search processes using GraphCodeBERT[13], and UniXcoder[12]. Our NMT models of learning ASTTrans Representation can boost the Mean Reciprocal Rank of these state-of-the-art code search processes by up to 3.08% and improve 23.08% of queries' results over the CAT benchmark.
Hung Phan, Ali Jannesari
CIKM2
2023 Performance Optimization using Multimodal Modeling and Heterogeneous GNN
abstract
Growing heterogeneity and configurability in HPC architectures has made auto-tuning applications and runtime parameters on these systems very complex. Users are presented with a multitude of options to configure parameters. In addition to application specific solutions, a common approach is to use general purpose search strategies, which often might not identify the best configurations or their time to convergence is a significant barrier. There is, thus, a need for a general purpose and efficient tuning approach that can be easily scaled and adapted to various tuning tasks. We propose a technique for tuning parallel code regions that is general enough to be adapted to multiple tasks. In this paper, we analyze IR-based programming models to make task-specific performance optimizations. To this end, we propose the Multimodal Graph Neural Network and Autoencoder (MGA) tuner, a multimodal deep learning based approach that adapts Heterogeneous Graph Neural Networks and Denoising Autoencoders for modeling IR-based code representations that serve as separate modalities. This approach is used as part of our pipeline to model a syntax, semantics, and structure-aware IR-based code representation for tuning parallel code regions/kernels. We extensively experiment on OpenMP and OpenCL code regions/kernels obtained from PolyBench, Rodinia, STREAM, DataRaceBench, AMD SDK, NPB, NVIDIA SDK, Parboil, SHOC, LULESH, XSBench, RSBench, miniFE, miniAMR, and Quicksilver benchmarks and applications. We apply our multimodal learning techniques to the tasks of (i) optimizing the number of threads, scheduling policy and chunk size in OpenMP loops and, (ii) identifying the best device for heterogeneous device mapping of OpenCL kernels. Our experiments show that this multimodal learning based approach outperforms the state-of-the-art in almost all experiments.
Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo César, Anna Sikora, Ali Jannesari
HPDC6
2023 Power Constrained Autotuning using Graph Neural Networks
abstract
Recent advances in multi and many-core processors have led to significant improvements in the performance of scientific computing applications. However, the addition of a large number of complex cores have also increased the overall power consumption, and power has become a first-order design constraint in modern processors. While we can limit power consumption by simply applying software-based power constraints, applying them blindly will lead to non-trivial performance degradation. To address the challenge of improving the performance, power, and energy efficiency of scientific applications on modern multi-core processors, we propose a novel Graph Neural Network based auto-tuning approach that (i) optimizes runtime performance at pre-defined power constraints, and (ii) simultaneously optimizes for runtime performance and energy efficiency by minimizing the energy-delay product. The key idea behind this approach lies in modeling parallel code regions as flow-aware code graphs to capture both semantic and structural code features. We demonstrate the efficacy of our approach by conducting an extensive evaluation on 30 benchmarks and proxy-/mini-applications with 68 OpenMP code regions. Our approach identifies OpenMP configurations at different power constraints that yield a geometric mean performance improvement of more than 25% and 13% over the default OpenMP configuration on a 32-core Skylake and a 16-core Haswell processor respectively. In addition, when we optimize for the energy-delay product, the OpenMP configurations selected by our auto-tuner demonstrate both performance improvement of 21% and 11% and energy reduction of 29% and 18% over the default OpenMP configuration at Thermal Design Power for the same Skylake and Haswell processors, respectively.
Akash Dutta, Jee Choi, Ali Jannesari
IPDPS3
2023 PERFOGRAPH: A Numerical Aware Program Graph Representation for Performance Optimization and Program Analysis
abstract
The remarkable growth and significant success of machine learning have expanded its applications into programming languages and program analysis. However, a key challenge in adopting the latest machine learning methods is the representation of programming languages which has a direct impact on the ability of machine learning methods to reason about programs. The absence of numerical awareness, aggregate data structure information, and improper way of presenting variables in previous representation works have limited their performances. To overcome the limitations and challenges of current program representations, we propose a novel graph-based program representation called PERFOGRAPH. PERFOGRAPH can capture numerical information and the aggregate data structure by introducing new nodes and edges. Furthermore, we propose an adapted embedding method to incorporate numerical awareness. These enhancements make PERFOGRAPH a highly flexible and scalable representation that can effectively capture programs' intricate dependencies and semantics. Consequently, it serves as a powerful tool for various applications such as program analysis, performance optimization, and parallelism discovery. Our experimental results demonstrate that PERFOGRAPH outperforms existing representations and sets new state-of-the-art results by reducing the error rate by 7.4% (AMD dataset) and 10% (NVIDIA dataset) in the well-known Device Mapping challenge. It also sets new state-of-the-art results in various performance optimization tasks like Parallelism Discovery and Numa and Prefetchers Configuration prediction.
Ali TehraniJamsaz, Quazi Ishtiaque Mahmud, Nesreen K. Ahmed, Ali Jannesari
NeurIPS5
2023 Reducing branch divergence to speed up parallel execution of unit testing on GPUs
Taghreed Bagies, Wei Le, Jeremy Sheaffer, Ali Jannesari
J. Supercomput.4
2022 Heterogeneous Graph Neural Networks for Software Effort Estimation
abstract
Background. Software effort can be measured by story point [35]. Story point estimation is important in software projects’ planning. Current approaches for automatically estimating story points focus on applying pre-trained embedding models and deep learning for text regression to solve this problem. These approaches require expensive embedding models and confront challenges that the sequence of text might not be an efficient representation for software issues which can be the combination of text and code.
Hung Phan, Ali Jannesari
ESEM2
2022 Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement Learning
abstract
Model compression is an essential technique for deploying deep neural networks (DNNs) on power and memory-constrained resources. However, existing model-compression methods often rely on human expertise and focus on parameters’ local importance, ignoring the rich topology information within DNNs. In this paper, we propose a novel multi-stage graph embedding technique based on graph neural networks (GNNs) to identify DNN topologies and use reinforcement learning (RL) to find a suitable compression policy. We performed resource-constrained (i.e., FLOPs) channel pruning and compared our approach with state-of-the-art model compression methods. We evaluated our method on various models from typical to mobile-friendly networks, such as ResNet family, VGG-16, MobileNet-v1/v2, and ShuffleNet. Results show that our method can achieve higher compression ratios with a minimal fine-tuning cost yet yields outstanding and competitive performance.
Sixing Yu, Arya Mazaheri, Ali Jannesari
ICML3
2022 Learning Intermediate Representations using Graph Neural Networks for NUMA and Prefetchers Optimization
abstract
There is a large space of NUMA and hardware prefetcher configurations that can significantly impact the performance of an application. Previous studies have demonstrated how a model can automatically select configurations based on the dynamic properties of the code to achieve speedups. This paper demonstrates how the static Intermediate Representation (IR) of the code can guide NUMA/prefetcher optimizations without the prohibitive cost of performance profiling. We propose a method to create a comprehensive dataset that includes a diverse set of intermediate representations along with optimum configurations. We then apply a graph neural network model in order to validate this dataset. We show that our static intermediate representation based model achieves 80 % of the performance gains provided by expensive dynamic performance profiling based strategies. We further develop a hybrid model that uses both static and dynamic information. Our hybrid model achieves the same gains as the dynamic models but at a reduced cost by only profiling 30 % of the programs.
Ali TehraniJamsaz, Mihail Popov, Akash Dutta, Emmanuelle Saillard, Ali Jannesari
IPDPS5
2022 SPATL: Salient Parameter Aggregation and Transfer Learning for Heterogeneous Federated Learning
abstract
Federated learning (FL) facilitates the training and deploying AI models on edge devices. Preserving user data privacy in FL introduces several challenges, including expensive communication costs, limited resources, and data heterogeneity. In this paper, we propose SPATL, an FL method that addresses these issues by: (a) introducing a salient parameter selection agent and communicating selected parameters only; (b) splitting a model into a shared encoder and a local predictor, and transferring its knowledge to heterogeneous clients via the locally customized predictor. Additionally, we leverage a gradient control mechanism to further speed up model convergence and increase robustness of training processes. Experiments demonstrate that SPATL reduces communication overhead, accelerates model inference, and enables stable training processes with better results compared to state-of-the-art methods. Our approach reduces communication cost by up to 86.45%, accelerates local inference by reducing up to 39.7% FLOPs on VGG-11, and requires 7.4× less communication overhead when training ResNet-20.11Code is available at: https://github.com/yusx-swapp/SPATL
Sixing Yu, Waqwoya Abebe, Ali Anwar 0001, Ali Jannesari
SC6
2021 Auto Graph Encoder-Decoder for Neural Network Pruning
Sixing Yu, Arya Mazaheri, Ali Jannesari
ICCV3
2021 A Learning-Based Scheduler for High Volume Processing in Data Warehouse Using Graph Neural Networks
Vivek Bengre, M. Reza HoseinyFarahabady, Mohammad Pivezhandi, Albert Y. Zomaya, Ali Jannesari
PDCAT5
2021 Building representative and balanced datasets of OpenMP parallel regions
abstract
Incorporating machine learning into automatic performance analysis and tuning tools is a promising path to tackle the increasing heterogeneity of current HPC applications. However, this introduces the need for generating balanced and representative datasets of parallel applications' executions. This work proposes a methodology for building datasets of OpenMP parallel code regions patterns. It allows for determining whether a given code region covers a unique part of the pattern input space not covered by the patterns already included in the dataset. The proposed methodology uses hardware performance counters to represent the execution of the region, which is referred to as the region signature for a given number of cores. Then, a complete representation of the region is built by joining the signatures for every different thread configuration in the system. Next, correlation analysis is performed between this representation and the representation of all the patterns already in the training set. Finally, if this correlation is below a given threshold, the region is considered to cover a unique part of the pattern input space and is subsequently added to the dataset. For validating this methodology, an example dataset, obtained from well known benchmarks, has been used to train a carefully designed neural network model to demonstrate that it is able to classify different patterns of OpenMP parallel regions.
Jordi Alcaraz, Steven Sleder, Ali TehraniJamsaz, Anna Sikora, Ali Jannesari, Joan Sorribes, Eduardo César
PDP5
2021 Reducing Energy in GPGPUs through Approximate Trivial Bypassing
abstract
General-purpose computing using graphics processing units (GPGPUs) is an attractive option for acceleration of applications with massively data-parallel tasks. While performance of modern GPGPUs is increasing rapidly, the power consumption of these devices is becoming a major concern. In particular, execution units and register file are among the top three most power-hungry components in GPGPUs. In this work, we exploit trivial instructions to reduce power consumption in GPGPUs. Trivial instructions are those instructions that do not need computations, i.e., multiplication by one. We found that, during the course of a program's execution, a GPGPU executes many trivial instructions. Execution of these instructions wastes power unnecessarily. In this work, we propose trivial bypassing which skips execution of trivial instructions and avoids unnecessary allocation of resources for trivial instructions. By power gating execution units and skipping trivial computing, trivial bypassing reduces both static and dynamic power. Also, trivial bypassing reduces dynamic energy of register file by avoiding access to register file for source and/or destination operands of trivial instructions. While trivial bypassing reduces energy of GPGPUs, it has detrimental impact on performance as a power-gated execution unit requires several cycles to resume its normal operation. Conventional warp schedulers are oblivious to the status of execution units. We propose a new warp scheduler that prioritizes warps based on availability of execution units. We also propose a set of new power management techniques to reduce performance penalty of power gating, further. To increase energy saving of trivial bypassing, we also propose approximating operands of instructions. We offer a set of new techniques to approximate both integer and floating-point instructions and increase the pool of trivial instructions. Our evaluations using a diverse set of benchmarks reveal that our proposed techniques are able to reduce energy of execution units by 11.2% and dynamic energy of register file by 12.2% with minimal performance and quality degradation.
Ehsan Atoofian, Zayan Shaikh, Ali Jannesari
ACM Trans. Embed. Comput. Syst.3
2020 Q-Flink: A QoS-Aware Controller for Apache Flink
abstract
Modern stream-data processing platforms are required to execute processing pipelines over high-volume, yet high-velocity, datasets under tight latency constraints. Apache Flink has emerged as an important new technology of large-scale platform that can distribute processing over a large number of computing nodes in a cluster (i.e., scale-out processing). Flink allows application developers to design and execute queries over continuous raw-inputs to analyze a large amount of streaming data in a parallel and distributed fashion. To increase the throughput of computing resources in stream processing platforms, a service provider might be tempted to use a consolidation strategy to pack as many processing applications as possible on the working nodes, with the hope of increasing the total revenue by improving the overall resource utilization. However, there is a hidden trap for achieving such a higher throughput solely by relying on an interference-oblivious consolidation strategy. In practice, collocated applications in a shared platform can fiercely compete with each others for obtaining the capacity of shared resources (e.g., cache and memory bandwidth) which in turn can lead to a severe performance degradation for all consolidated workloads.This paper addresses the shared resource contention problem associated with the auto-resource controlling mechanism of Apache Flink engine running across a distributed cluster. A controlling strategy is proposed to handle scenarios in which stream processing applications may have different quality of service (QoS) requirements while the resource interference is considered as the key performance-limiting parameter. The performance evaluation is carried out by comparing the proposed controller with the default Flink resource allocation strategy in a testbed cluster with total 32 Intel Xeon cores under different workload traffic with up to 4000 streaming applications chosen from various benchmarking tools. Experimental results demonstrate that the proposed controller can successfully decrease the average latency of high priority applications by 223% during the burst traffic while maintaining the requested QoS enforcement levels.
M. Reza HoseinyFarahabady, Ali Jannesari, Javid Taheri, Wei Bao 0001, Albert Y. Zomaya, Zahir Tari
CCGRID2
2020 Skipping Non-essential Instructions Makes Data-Dependence Profiling Faster
Nicolas Morew, Mohammad Norouzi Arab, Ali Jannesari, Felix Wolf 0001
Euro-Par3
2020 Accelerating winograd convolutions using symbolic computation and meta-programming
abstract
Convolution operations are essential constituents of convolutional neural networks. Their efficient and performance-portable implementation demands tremendous programming effort and fine-tuning. Winograd's minimal filtering algorithm is a well-known method to reduce the computational complexity of convolution operations. Unfortunately, existing implementations of this algorithm are either vendor-specific or hard-coded to support a small subset of convolutions, thus limiting their versatility and performance portability. In this paper, we propose a novel method to optimize Winograd convolutions based on symbolic computation. Taking advantage meta-programming and auto-tuning, we further introduce a system to automate the generation of efficient and portable Winograd convolution code for various GPUs. We show that our optimization technique can effectively exploit repetitive patterns, enabling us to reduce the number of arithmetic operations by up to 62% without compromising numerical stability. Moreover, we demonstrate in experiments that we can generate efficient kernels with runtimes close to deep-learning libraries, requiring only a minimum of programming effort, which confirms the performance portability of our approach.
Arya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf 0001, Ali Jannesari
EuroSys5
2019 Accelerating Data-Dependence Profiling with Static Hints
Mohammad Norouzi Arab, Qamar Ilias, Ali Jannesari, Felix Wolf 0001
Euro-Par3
2019 Enhancing the Programmability and Performance Portability of GPU Tensor Operations
Arya Mazaheri, Johannes Schulte, Matthew W. Moskewicz, Felix Wolf 0001, Ali Jannesari
Euro-Par5
2019 Automatic construct selection and variable classification in OpenMP
abstract
A major task of parallelization with OpenMP is to decide where in a program to insert which OpenMP construct such that speedup is maximized and correctness is preserved. Another challenge is the classification of variables that appear in a construct according to their data-sharing semantics. Manual classification is tedious and error prone. Moreover, the choice of the data-sharing attribute can significantly affect performance. Grounded on the notion of parallel design patterns, we propose a method that identifies code regions to parallelize and selects appropriate OpenMP constructs for them. Also, we classify variables in the chosen constructs by analyzing data dependences that have been dynamically extracted from the program. Using our approach, we created OpenMP versions of 49 sequential benchmarks and compared them with the code produced by three state-of-the-art parallelization tools: Our codes are faster in most cases with average speedups relative to any of the three ranging from 1.8 to 2.7. Additionally, we automatically reclassified variables of OpenMP programs parallelized manually or with the help of these tools, improving their execution time by up to 29%.
Mohammad Norouzi Arab, Felix Wolf 0001, Ali Jannesari
ICS3
2019 Dynamic Control of CPU Cap Allocations in Stream Processing and Data-Flow Platforms
abstract
This paper focuses on Timely dataflow programming model for processing streams of data. We propose a technique to define CPU resource allocation (i.e., CPU capping) with the goal to improve response time latency in such type of applications with different quality of service (QoS) level, as they are concurrently running in a shared multi-core computing system with unknown and volatile demand. The proposed solution predicts the expected performance of the underlying platform using an online approach based on queuing theory and adjusts the corrections required in CPU allocation to achieve the most optimized performance. The experimental results confirms that measured performance of the proposed model is highly accurate while it takes into account the percentiles on the QoS metrics. The theoretical model used for elastic allocation of CPU share in the target platform takes advantage of design principals in model predictive control theory and dynamic programming to solve an optimization problem. While the prediction module in the proposed algorithm tries to predict the temporal changes in the arrival rate of each data flow, the optimization module uses a system model to estimate the interference among collocated applications by continuously monitoring the available CPU utilization in individual nodes along with the number of outstanding messages in every intermediate buffer of all TDF applications. The optimization module eventually performs a cost-benefit analysis to mitigate the total amount of QoS violation incidents by assigning the limited CPU shares among collocated applications. The proposed algorithm is robust (i.e., its worst-case output is guaranteed for arbitrarily volatile incoming demand coming from different data streams), and if the demand volatility is not large, the output is optimal, too. Its implementation is done using the TDF framework in Rust for distributed and shared memory architectures. The experimental results show that the proposed algorithm reduces the average and p99 latency of delay-sensitive applications by 21% and 31.8%, respectively, while can reduce the amount of QoS violation incidents by 98% on average.
M. Reza HoseinyFarahabady, Ali Jannesari, Zahir Tari, Javid Taheri, Albert Y. Zomaya
NCA2
2019 Real-Time Stream Data Processing at Scale
abstract
A typical scenario in a stream data-flow processing engine is that users submit continues queries in order to receive the computational result once a new stream of data arrives. The focus of the paper is to design a dynamic CPU cap controller for stream data-flow applications with real-time constraints, in which the result of computations must be available within a short time period, specified by the user, once a recent update in the input data occurs. It is common that the stream data-flow processing engine is deployed over a cluster of dedicated or virtualized server nodes, e.g., Cloud or Edge platform, to achieve a faster data processing. However, the attributes of incoming stream data-flow might fluctuate in an irregular way. To effectively cope with such unpredictable conditions, the underlying resource manager needs to be equipped with a dynamic resource provisioning mechanism to ensure the real-time requirements of different applications. The proposed solution uses control theory principals to achieve a good utilization of computing resources and a reduced average response time. The proposed algorithm dynamically adjusts the required quality of service (QoS) in an environment when multiple stream & data-flow processing applications concurrently run with unknown and volatile workloads. Our study confirms that such a unpredictable demand can negatively degrade the system performance, mainly due to adverse interference in the utilization of shared resources. Unlike prior research studies which assumes a static or zero correlation among the performance variability among consolidated applications, we presume the prevalence of shared-resource interference among collocated applications as a key performance-limiting parameter and confront it in scenarios where several applications have different QoS requirements with unpredictable workload demands. We design a low-overhead controller to achieve two natural optimization objectives of minimizing QoS violation amount and maximizing the average CPU utilization. The algorithm takes advantage of design principals in model predictive control theory for elastic allocation of CPU share. The experimental results confirm that there is a strong correlation in performance degradation among consolidation strategies and the system utilization for obtaining the capacity of shared resources in a non-cooperative manner. The results confirm that the proposed solution can reduce the average latency of delay-sensitive applications by 17% comparing to the results of a well established heuristic called Class-Based Weighted Fair Queuing (CFWFQ). At the same time, the proposed solution can prevent the QoS violation incidents by 62%.
M. Reza HoseinyFarahabady, Ali Jannesari, Wei Bao 0001, Zahir Tari, Albert Y. Zomaya
PDCAT2
2019 Dissecting sequential programs for parallelization - An approach based on computational units
abstract
Summary When trying to parallelize a sequential program, programmers routinely struggle during the first step: finding out which code sections can be made to run in parallel. While identifying such code sections, most of the current parallelism discovery techniques focus on specific language constructs. In contrast, we propose to concentrate on the computations performed by a program. In our approach, a program is treated as a collection of computations communicating with one another using a number of variables. Each computation is represented as a computational unit (CU). A CU contains the inputs and outputs of a computation, and the three phases of a computation are read, compute, and write. Based on the notion of CU, which ensures that the read phase executes before the write phase, we present a unified framework to identify both loop parallelism and task parallelism in sequential programs. We conducted a range of experiments on 23 applications from four different benchmark suites. Our approach accurately identified the parallelization opportunities in benchmark applications based on comparison with their parallel versions. We have also parallelized the opportunities identified by our approach that were not implemented in the parallel versions of the benchmarks and reported the speedup.
Rohit Atre, Zia Ul Huda, Felix Wolf 0001, Ali Jannesari
Concurr. Comput. Pract. Exp.4
2019 The Art of Getting Deep Neural Networks in Shape
abstract
Training a deep neural network (DNN) involves selecting a set of hyperparameters that define the network topology and influence the accuracy of the resulting network. Often, the goal is to maximize prediction accuracy on a given dataset. However, non-functional requirements of the trained network -- such as inference speed, size, and energy consumption -- can be very important as well. In this article, we aim to automate the process of selecting an appropriate DNN topology that fulfills both functional and non-functional requirements of the application. Specifically, we focus on tuning two important hyperparameters, depth and width, which together define the shape of the resulting network and directly affect its accuracy, speed, size, and energy consumption. To reduce the time needed to search the design space, we train a fraction of DNNs and build a model to predict the performances of the remaining ones. We are able to produce tuned ResNets, which are up to 4.22 times faster than original depth-scaled ResNets on a batch of 128 images while matching their accuracy.
Rahim Mammadli, Felix Wolf 0001, Ali Jannesari
ACM Trans. Archit. Code Optim.3
2018 Unveiling Thread Communication Bottlenecks Using Hardware-Independent Metrics
abstract
A critical factor for developing robust shared-memory applications is the efficient use of the cache and the communication between threads. Inappropriate data structures, algorithm design, and inefficient thread affinity may result in superfluous communication between threads/cores and severe performance problems. For this reason, state-of-the-art profiling tools focus on thread communication and behavior to present different metrics that enable programmers to write cache-friendly programs. The data shared between a pair of threads should be reused with a reasonable distance to preserve data locality. However, existing tools do not take into account the locality of communication events and mainly focus on analyzing the amount of communication instead. In this paper, we introduce a new method to analyze performance and communication bottlenecks that arise from data-access patterns and thread interactions of each code region. We propose new hardware-independent metrics to characterize thread communication and provide suggestions for applying appropriate optimizations on a specific code region. We evaluated our approach on the SPLASH and Rodinia benchmark suites. Experimental results validate the effectiveness of our approach by finding communication locality issues due to inefficient data structures and/or poor algorithm implementations. By applying the suggested optimizations, we improved the performance in Rodinia benchmarks by up to 56%. Furthermore, by varying the input size we demonstrated the ability of our method to assess the cache usage and scalability of a given application in terms of its inherent communication.
Arya Mazaheri, Felix Wolf 0001, Ali Jannesari
ICPP3
2018 Improving performance of transactional memory through machine learning
abstract
Summary Transactional memory (TM) is a programming paradigm that facilitates parallel programming for multi‐core processors. In the last few years, some chip manufacturers provided hardware support for TM to reduce runtime overhead of Software Transactional Memory (STM). In this work, we offer two optimization techniques for TMs. The first technique focuses on Restricted Transactional Memory (RTM) in Intel's Haswell processor and shows that while in some applications, RTM improves performance over STM, in some others, it falls behind STM. We exploit this variability and propose an adaptive technique that switches between RTM and STM, statically. The second technique focuses on the overhead of TM and enhances the speed of the adaptive system. In particular, we focus on the size of transactions and improve performance by changing the transaction size. Optimizing the transaction size manually is a time‐consuming process and requires significant software engineering effort. We use a combination of Linear Regression (LR) and decision tree to decide on the transaction size, automatically. We evaluate our optimization techniques using a set of benchmarks from NAS, DiscoPoP, and STAMP benchmark suites. Our experimental results reveal that our optimization techniques are able to improve the performance of TM programs by 9% and energy‐delay by 15%, on average.
Yang Xiao 0001, Thireshan Jeyakumaran, Ehsan Atoofian, Ali Jannesari
Concurr. Comput. Pract. Exp.4
2017 Brief Announcement: Meeting the Challenges of Parallelizing Sequential Programs
abstract
Discovering which code sections in a sequential program can be made to run in parallel is the first step in parallelizing it, and programmers routinely struggle in this step. Most of the current parallelism discovery techniques focus on specific language constructs while trying to identify such code sections. In contrast, we propose to concentrate on the computations performed by a program. In our approach, a program is treated as a collection of computations communicating with one another using a number of variables. Each computation is represented as a Computational Unit (CU). A CU contains the inputs and outputs of a computation, and the three phases of a computation: read, compute, and write. Based on the notion of CU, We present a unified framework to identify both loop and task parallelism in sequential programs.
Rohit Atre, Ali Jannesari, Felix Wolf 0001
SPAA2
2017 Editorial of special issue on Software Engineering for Parallel Systems
Ali Jannesari, Felix Wolf 0001, Walter F. Tichy
J. Syst. Softw.1
2016 Automatic Parallel Pattern Detection in the Algorithm Structure Design Space
abstract
Parallel design patterns have been developed to help programmers efficiently design and implement parallel applications. However, identifying a suitable parallel pattern for a specific code region in a sequential application is a difficult task. Transforming an application according to support structures applicable to these parallel patterns is also very challenging. In this paper, we present a novel approach to automatically find parallel patterns in the algorithm structure design space of sequential applications. In our approach, we classify code blocks in a region according to the appropriate supportstructure of the detected pattern. This classification eases the transformation of a sequential application into its parallel version. Weevaluated our approach on 17 applications from four different benchmark suites. Our method identified suitable algorithm structure patterns in the sequential applications. We confirmed our results by comparing them with the existing parallel versions of these applications. We also implemented the patterns we detected in cases in which parallel implementations were not available and achieved speedups of up to 14x.
Zia Ul Huda, Rohit Atre, Ali Jannesari, Felix Wolf 0001
IPDPS3
2016 Improving Performance of Transactional Applications through Adaptive Transactional Memory
abstract
Transactional memory (TM) has become progressively widespread especially with hardware transactional memory implementation becoming increasingly available. In this paper, we focus on Restricted Transactional Memory (RTM) in Intel's Haswell processor and show that performance of RTM varies across applications. While RTM enhances performance of some applications relative to software transactional memory (STM), in some others, it degrades performance. We exploit this variability and present an adaptive system which is a static approach that switches between HTM and STM in transaction granularity. By incorporating a decision tree prediction module, we are able to predict the optimum TM system for a given transaction based on its characteristics. Our adaptive system supports both HTM and STM with the aim of increasing an application's performance. We show that our adaptive system has an average overall speedup of 20.82% over both TM systems.
Thireshan Jeyakumaran, Ehsan Atoofian, Yang Xiao 0001, Zhen Li 0005, Ali Jannesari
PDP5
2016 Unveiling parallelization opportunities in sequential programs
Zhen Li 0005, Rohit Atre, Zia Ul Huda, Ali Jannesari, Felix Wolf 0001
J. Syst. Softw.4
2015 Fast Data-Dependence Profiling by Skipping Repeatedly Executed Memory Operations
Zhen Li 0005, Michael Beaumont, Ali Jannesari, Felix Wolf 0001
ICA3PP (4)3
2015 Beyond Data Parallelism: Identifying Parallel Tasks in Sequential Programs
Zhen Li 0005, Bo Zhao 0019, Ali Jannesari, Felix Wolf 0001
ICA3PP (4)3
2015 Automatic Optimization of Software Transactional Memory Through Linear Regression and Decision Tree
Yang Xiao 0001, Zhen Li 0005, Ehsan Atoofian, Ali Jannesari
ICA3PP (4)4
2015 Characterizing Loop-Level Communication Patterns in Shared Memory
abstract
Communication patterns extracted from parallel programs can provide a valuable source of information for parallel pattern detection, application auto-tuning, and runtime workload scheduling on heterogeneous systems. Once identified, such patterns can help find the most promising optimizations. Communication patterns can be detected using different methods, including sandbox simulation, memory profiling, and hardware counter analysis. However, these analyses usually suffer from high runtime and memory overhead, necessitating a trade off between accuracy and resource consumption. More importantly, none of the existing methods exploit fine-grained communication patterns on the level of individual code regions. In this paper, we present an efficient tool based on Disco PoP profiler that characterizes the communication pattern of every hotspot in a shared-memory application. With the aid of static and dynamic code analysis, it produces a nested structure of communication patterns based on program's loops. By employing asymmetric signature memory, the runtime overhead is around 225× while the required amount of memory remains fixed. In comparison with other profilers, the proposed method is efficient enough to be used with real world applications.
Arya Mazaheri, Ali Jannesari, Abdolreza Mirzaei, Felix Wolf 0001
ICPP2
2015 An Efficient Data-Dependence Profiler for Sequential and Parallel Programs
abstract
Extracting data dependences from programs serves as the foundation of many program analysis and transformation methods, including automatic parallelization, runtime scheduling, and performance tuning. To obtain data dependences, more and more related tools are adopting profiling approaches because they can track dynamically allocated memory, pointers, and array indices. However, dependence profiling suffers from high runtime and space overhead. To lower the overhead, earlier dependence profiling techniques exploit features of the specific program analyses they are designed for. As a result, every program analysis tool in need of data-dependence information requires its own customized profiler. In this paper, we present an efficient and at the same time generic data-dependence profiler that can be used as a uniform basis for different dependence-based program analyses. Its lock-free parallel design reduces the runtime overhead to around 86× on average. Moreover, signature-based memory management adjusts space requirements to practical needs. Finally, to support analyses and tuning approaches for parallel programs such as communication pattern detection, our profiler produces detailed dependence records not only for sequential but also for multi-threaded code.
Zhen Li 0005, Ali Jannesari, Felix Wolf 0001
IPDPS2
2015 Resource and application-aware resource discovery in computing environments
Mohammad Norouzi Arab, Ali Jannesari
J. Supercomput.2
2014 Using Template Matching to Infer Parallel Design Patterns
abstract
The triumphant spread of multicore processors over the past decade increases the pressure on software developers to exploit the growing amount of parallelism available in the hardware. However, writing parallel programs is generally challenging. For sequential programs, the formulation of design patterns marked a turning point in software development, boosting programmer productivity and leading to more reusable and maintainable code. While the literature is now also reporting a rising number of parallel design patterns, programmers confronted with the task of parallelizing an existing sequential program still struggle with the question of which parallel pattern to apply where in their code. In this article, we show how template matching, a technique traditionally used in the discovery of sequential design patterns, can also be used to support parallelization decisions. After looking for matches in a previously extracted dynamic dependence graph, we classify code blocks of the input program according to the structure of the parallel patterns we find. Based on this information, the programmer can easily implement the detected pattern and create a parallel version of his or her program. We tested our approach with six programs, in which we successfully detected pipeline and do-all patterns.
Zia Ul Huda, Ali Jannesari, Felix Wolf 0001
ACM Trans. Archit. Code Optim.2
2014 Library-Independent Data Race Detection
abstract
Data races are a common problem on shared-memory parallel computers, including multicores. Analysis programs called race detectors help find and eliminate them. However, current race detectors are geared for specific concurrency libraries. When programmers use libraries unknown to a given detector, the detector becomes useless or requires extensive reprogramming. We introduce a new synchronization detection mechanism that is independent of concurrency libraries. It dynamically detects synchronization constructs based on a characteristic code pattern. The approach is non-intrusive and applicable to various concurrency libraries. Experimental results confirm that the approach identifies synchronizations and detects data races regardless of the concurrency libraries involved. With this mechanism, race detectors can be written once and need not be adapted to particular libraries.
Ali Jannesari, Walter F. Tichy
IEEE Trans. Parallel Distributed Syst.1
2013 Predicting Parallelization of Sequential Programs Using Supervised Learning
abstract
We investigate an automatic method for classifying which regions of sequential programs could be parallelized, using dynamic features of the code collected at runtime. We train a supervised learning algorithm on versions of the NAS Parallel Benchmark (NPB) code hand-annotated with OpenMP parallelization directives in order to approximate the parallelization that might be produced by a human expert. A model comparison shows that support vector machines and decision trees have comparable performance on this classification problem, but boosting using AdaBoost is able to increase the performance of the decision trees. We further analyze the relative importance of the collected program features and demonstrate that within-loop instruction counts provide the greatest contribution to decision tree error reduction, with dependency graph features of secondary importance.
Daniel Fried, Zhen Li 0005, Ali Jannesari, Felix Wolf 0001
ICMLA (2)3
2013 Detecting Correlation Violations and Data Races by Inferring Non-deterministic Reads
abstract
With the introduction of multicore systems and parallel programs concurrency bugs have become more common. A notorious class of these bugs are data races that violate correlations between variables. This happens, for example, when the programmer does not update correlated variables atomically, which is needed to maintain their semantic relationship. The detection of such races is challenging because correlations among variables usually escape traditional race detectors which are oblivious of semantic relationships. In this paper, we present an effective method for dynamically identifying correlated variables together with a race detector based on the notion of non-deterministic reads that identifies malicious data races on correlated variables. In eight programs and 190 micro benchmarks, we found more than 100 races that were overlooked by other race detectors. Furthermore, we identified about 300 variable correlations which were violated by these races.
Ali Jannesari, Nico Koprowski, Jochen Schimmel, Felix Wolf 0001, Walter F. Tichy
ICPADS1
2013 Discovery of Potential Parallelism in Sequential Programs
abstract
Although multicore CPUs are dominating the market of desktops and servers, writing programs that utilize the available hardware parallelism on these architectures still remains a challenge. In this paper, we present a dynamic approach for automatically identifying potential parallelism in sequential programs. Our method is based on the notion of computational units, which are small sections of code following the read-compute-write pattern that can form the atoms of concurrent scheduling. In contrast to earlier approaches, our method can identify parallelism between code sections of arbitrary granularity and does not rely on a predefined notion of language constructs subject to parallelization. Experimental results show that reasonable speedups can be achieved by parallelizing sequential programs manually according to our findings. By comparing our findings to known parallel implementations of sequential programs, we demonstrate that we are able to detect the most important code locations to be parallelized.
Zhen Li 0005, Ali Jannesari, Felix Wolf 0001
ICPP2
2011 Dynamic Data Race Detection for Correlated Variables
Ali Jannesari, Markus Westphal-Furuya, Walter F. Tichy
ICA3PP (1)1
2010 Identifying ad-hoc synchronization for enhanced race detection
abstract
Parallel programs contain a surprising number of ad-hoc synchronization operations. Ad-hoc synchronization operations are loops that busy-wait on condition variables. Current race detectors produce unnecessary warnings (false positives) when ad-hoc synchronization is used. False positives are also generated when programmers use synchronization primitives that are unknown to race detectors, for instance when programmers switch libraries. These shortcomings may result in an overwhelming number of false positives, dissuading programmers from using race detectors. This paper shows that ad-hoc synchronization operations can be detected automatically. The method requires no user intervention such as annotations and has been implemented in the race detector Helgrind+. Evaluation results on various benchmarks confirm that Helgrind+is aware of all synchronizations in programs, reliably reports true races, and produces few false alarms. A surprising result is that with the new technique, Helgrind+can analyze synchronization libraries, so special knowledge about these libraries is not needed in the detector.
Ali Jannesari, Walter F. Tichy
IPDPS1
2009 Helgrind+: An efficient dynamic race detector
abstract
Finding synchronization defects is difficult due to non-deterministic orderings of parallel threads. Current tools for detecting synchronization defects tend to miss many data races or produce an overwhelming number of false alarms. In this paper, we describe Helgrind+, a dynamic race detection tool that incorporates correct handling of condition variables and a combination of the lockset algorithm and happens-before relation. We compare our techniques with Intel Thread Checker and the original Helgrind tool on two substantial benchmark suites. Helgrind+ reduces the number of both false negatives (missed races) and false positives. The additional accuracy incurs almost no performance overhead.
Ali Jannesari, Kaibin Bao, Victor Pankratius, Walter F. Tichy
IPDPS1