Akash Dutta

dblp:272/6678 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 EEG-Based Hybrid Emotion Recognition Model with Statistical-Wavelet Features and Modality-Agnostic Loss
Asfak Ali, Jotiraditya Banerjee, Debam Saha, Akash Dutta, Friedhelm Schwenker, Ram Sarkar
EANN (1)4
2025 HHOTuner: Efficient Performance Tuning with Harris Hawks Optimization
abstract
As we enter the Post-Exascale Computing era, parallel programs are becoming more ubiquitous. Increased parallelization efforts and larger, more complex systems have led to numerous ways of tweaking performance of such applications. Existing tuners often produce very good results, but are often task specific, or have high overheads. Our aim is to enable and adapt optimization techniques seen in nature for auto-tuning HPC search spaces. We believe that such an approach can outperform the auto-tuning abilities of state-of-the-art tuners while significantly reducing associated overheads. To this end, we propose the HHOTuner (Harris Hawks Optimization Tuner): a nature-inspired meta-heuristic, swarm-based optimization technique for auto-tuning real-world HPC search spaces. Named after the Harris hawks, the HHOTuner employs and adapts the real-world predatory behavior of Harris hawks as algorithms for tuning real-world HPC applications and workloads. The proposed auto-tuner is general purpose and is designed to work well with user-defined search spaces. We have evaluated HHOTuner on several workloads such as HPC applications, graph neural network epoch training, and tensor program generation by an auto-tuning compiler. HHOTuner improves, sometimes significantly, the performance of these workloads. Its performance is better than state-of-the-art auto-tuners, sometimes being an order of magnitude ahead for a few cases. HHOTuner’s tuning overhead is much lower (up to 7.2 × and 32.9 ×) than the other auto-tuners evaluated in this paper, while maintaining or improving the quality of auto-tuning results.
Akash Dutta, Ali Jannesari
ICPP1
2024 MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
abstract
One of the primary areas of interest in High Performance Computing is the improvement of performance of parallel workloads. Nowadays, compilable source code-based optimization tasks that employ deep learning often exploit LLVM Intermediate Representations (IRs) for extracting features from source code. Most such works target specific tasks, or are designed with a pre-defined set of heuristics. So far, pre-trained models are rare in this domain, but the possibilities have been widely discussed. Especially approaches mimicking large-language models (LLMs) have been proposed. But these have prohibitively large training costs. In this paper, we propose MIREncoder, a Multi-modal IR-based Auto-Encoder that can be pre-trained to generate a learned embedding space to be used for downstream tasks by machine learning-based approaches. A multi-modal approach enables us to better extract features from compilable programs. It allows us to better model code syntax, semantics and structure. For code-based performance optimizations, these features are very important while making optimization decisions. A pre-trained model/embedding implicitly enables the usage of transfer learning, and helps move away from task-specific trained models. Additionally, a pre-trained model used for downstream performance optimization should itself have reduced overhead, and be easily usable. These considerations have led us to propose a modeling approach that i) understands code semantics and structure, ii) enables use of transfer learning, and iii) is small and simple enough to be easily re-purposed or reused even with low resource availability. Our evaluations will show that our proposed approach can outperform the state of the art while reducing overhead.
Akash Dutta, Ali Jannesari
PACT1
2024 Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
abstract
Abstract Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Akash Dutta, Ali Jannesari, Jordi Alcaraz
Euro-Par (1)5
2024 Static Generation of Efficient OpenMP Offload Data Mappings
abstract
Increasing heterogeneity in HPC architectures and compiler advancements have led to OpenMP being frequently used to enable computations on heterogeneous devices. However, the efficient movement of data on heterogeneous computing platforms is crucial for achieving high utilization. Programmers must explicitly map data between the host and connected accelerator devices to achieve efficient data movement. Ensuring efficient data transfer requires programmers to reason about complex data flow. This can be a laborious and error-prone process since the programmer must keep a mental model of data validity and lifetime spanning multiple data environments. We present a static analysis tool, OMPDart (OpenMP Data Reduction Tool), for OpenMP programs that models data dependencies between host and device regions and applies source code transformations to achieve efficient data transfer. Our evaluations on nine HPC benchmarks demonstrate that OMPDart is capable of generating effective data mapping constructs that substantially reduce data transfer between host and device.
Luke J. Marzen, Akash Dutta, Ali Jannesari
SC2
2023 Performance Optimization using Multimodal Modeling and Heterogeneous GNN
abstract
Growing heterogeneity and configurability in HPC architectures has made auto-tuning applications and runtime parameters on these systems very complex. Users are presented with a multitude of options to configure parameters. In addition to application specific solutions, a common approach is to use general purpose search strategies, which often might not identify the best configurations or their time to convergence is a significant barrier. There is, thus, a need for a general purpose and efficient tuning approach that can be easily scaled and adapted to various tuning tasks. We propose a technique for tuning parallel code regions that is general enough to be adapted to multiple tasks. In this paper, we analyze IR-based programming models to make task-specific performance optimizations. To this end, we propose the Multimodal Graph Neural Network and Autoencoder (MGA) tuner, a multimodal deep learning based approach that adapts Heterogeneous Graph Neural Networks and Denoising Autoencoders for modeling IR-based code representations that serve as separate modalities. This approach is used as part of our pipeline to model a syntax, semantics, and structure-aware IR-based code representation for tuning parallel code regions/kernels. We extensively experiment on OpenMP and OpenCL code regions/kernels obtained from PolyBench, Rodinia, STREAM, DataRaceBench, AMD SDK, NPB, NVIDIA SDK, Parboil, SHOC, LULESH, XSBench, RSBench, miniFE, miniAMR, and Quicksilver benchmarks and applications. We apply our multimodal learning techniques to the tasks of (i) optimizing the number of threads, scheduling policy and chunk size in OpenMP loops and, (ii) identifying the best device for heterogeneous device mapping of OpenCL kernels. Our experiments show that this multimodal learning based approach outperforms the state-of-the-art in almost all experiments.
Akash Dutta, Jordi Alcaraz, Ali TehraniJamsaz, Eduardo César, Anna Sikora, Ali Jannesari
HPDC1
2023 Power Constrained Autotuning using Graph Neural Networks
abstract
Recent advances in multi and many-core processors have led to significant improvements in the performance of scientific computing applications. However, the addition of a large number of complex cores have also increased the overall power consumption, and power has become a first-order design constraint in modern processors. While we can limit power consumption by simply applying software-based power constraints, applying them blindly will lead to non-trivial performance degradation. To address the challenge of improving the performance, power, and energy efficiency of scientific applications on modern multi-core processors, we propose a novel Graph Neural Network based auto-tuning approach that (i) optimizes runtime performance at pre-defined power constraints, and (ii) simultaneously optimizes for runtime performance and energy efficiency by minimizing the energy-delay product. The key idea behind this approach lies in modeling parallel code regions as flow-aware code graphs to capture both semantic and structural code features. We demonstrate the efficacy of our approach by conducting an extensive evaluation on 30 benchmarks and proxy-/mini-applications with 68 OpenMP code regions. Our approach identifies OpenMP configurations at different power constraints that yield a geometric mean performance improvement of more than 25% and 13% over the default OpenMP configuration on a 32-core Skylake and a 16-core Haswell processor respectively. In addition, when we optimize for the energy-delay product, the OpenMP configurations selected by our auto-tuner demonstrate both performance improvement of 21% and 11% and energy reduction of 29% and 18% over the default OpenMP configuration at Thermal Design Power for the same Skylake and Haswell processors, respectively.
Akash Dutta, Jee Choi, Ali Jannesari
IPDPS1
2022 Learning Intermediate Representations using Graph Neural Networks for NUMA and Prefetchers Optimization
abstract
There is a large space of NUMA and hardware prefetcher configurations that can significantly impact the performance of an application. Previous studies have demonstrated how a model can automatically select configurations based on the dynamic properties of the code to achieve speedups. This paper demonstrates how the static Intermediate Representation (IR) of the code can guide NUMA/prefetcher optimizations without the prohibitive cost of performance profiling. We propose a method to create a comprehensive dataset that includes a diverse set of intermediate representations along with optimum configurations. We then apply a graph neural network model in order to validate this dataset. We show that our static intermediate representation based model achieves 80 % of the performance gains provided by expensive dynamic performance profiling based strategies. We further develop a hybrid model that uses both static and dynamic information. Our hybrid model achieves the same gains as the dynamic models but at a reduced cost by only profiling 30 % of the programs.
Ali TehraniJamsaz, Mihail Popov, Akash Dutta, Emmanuelle Saillard, Ali Jannesari
IPDPS3