EDBT 2026 Demo / reviewers in the wild / expert
Preeti Malakar
dblp:05/9135
· DBLP profile ↗
18ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-7038-8712ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 8 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Communication-Balanced Job Allocation Using SLURM
Gagandeep Mangat, Preeti Malakar |
JSSPP | 2 |
| 2025 | Fine-grained Communication Phase based Analytical Performance Modeling and AnalysisabstractHigh-performance computing is essential for scientific innovation. With the advent of exascale computing and the growing scale of scientific workloads, novel tools and methodologies are extremely important to analyze and model the performance of large-scale scientific applications. Existing profiling and tracing tools have certain limitations: profiles do not allow fine-grained performance modeling, whereas trace-based simulation or performance modeling is expensive and often infeasible for large applications. In this work, we propose fine-grained communication phases as a level of abstraction for analyzing and modeling application performance. Given the iterative nature of HPC applications, we propose a methodology to automatically group MPI communication events into phases and identify the unique and repeating phases. Our approach enables modeling only the unique phases, in contrast to a complete trace simulation. Further, we address the limitation of existing analytical communication models to model runtime delays in communication time. We propose a new delay-aware communication model (DACM) that achieves a best-case prediction error of 10.11% across six diverse HPC benchmarks and applications, in contrast to 53% using existing analytical models. Vishal Deka, Preeti Malakar |
SBAC-PAD | 2 |
| 2024 | Evaluating Active-learning Based Performance Prediction of Parallel ApplicationsabstractMany scientific applications use Message Passing Interface for distributed-memory parallelism. These applications simulate complex physical phenomena and are routinely executed on supercomputers and high performance compute clusters. The applications may run for several hours or days and there may be long queue waiting times on the supercomputers. Often the application developers are unable to correctly estimate the runtime of their application for a new configuration and thus may overestimate the runtime requirements. This may increase the queue wait times along with sub-optimal job scheduling decisions. Thus accurate prediction of the performance of an application (performance modeling) is helpful. We have developed statistical models for performance prediction of MPI applications based on active learning algorithms and are able to achieve low Median Absolute Percentage Error (MdAPE) in many cases. We used five scientific applications and benchmarks – HPCG, LULESH, miniAMR, miniMD, and miniFE. We predict the execution times of these applications using various data sizes on up to 512 cores of the PARAM Sanganak supercomputer at IIT Kanpur. The statistical models gave MdAPE of 10.6, 3.5, 8.21, 12.62, and 32.25% for HPCG, LULESH, miniAMR, miniMD, and miniFE respectively. Shivam Aggarwal, Preeti Malakar |
e-Science | 2 |
| 2023 | CCGRID 2023: A Holistic Approach to Inclusion and Belongingabstract“CCGRID will act with responsibility as its primary consideration; with equity, diversity, and inclusion as its central goals.” from the CCGRID 2023 web site [1] Beth Plale, Preeti Malakar, Meenakshi D'Souza, Hemangee K. Kapoor, Yogesh L. Simmhan, Ilkay Altintas, S. Manohar 0001 |
CCGrid | 2 |
| 2022 | A Deep Learning-Based In Situ Analysis Framework for Tropical Cyclogenesis PredictionabstractTropical cyclone is one of the most violent natural disasters causing massive devastation. Accurate forecasting of cyclones with high lead times is an important problem. We propose a framework to predict tropical cyclogenesis (i.e. cyclone formation). This framework executes along with a parallel weather simulation model (WRF) and analyzes the simulation output as soon as they are generated. Our framework has two major components – a trigger function and a deep predictive model. The trigger function acts as a basic filter to identify cyclones from non-cyclones. The proposed deep learning model is based on convolutional neural networks (CNNs). The best track data from Indian Meteorological Department (IMD) is used as a reference for labeling data points into disturbances and tropical cyclones. The framework achieves a probability of detection (POD) value of approximately 95% with a false alarm ratio (FAR) of 21.69% overall. The predictions made by the framework have a lead time of up to 150 hours from the time that a disturbance transforms into a tropical cyclone. Abir Mukherjee, Preeti Malakar |
HIPC | 2 |
| 2021 | An Integrated Job Monitor, Analyzer and PredictorabstractHigh performance computing systems are used for compute-intensive jobs by multiple users. The users submit jobs to batch queues where the jobs are queued for an unknown amount of time until the required resources are available. A large amount of data (submit time, start time, end time, nodes allocated) is collected about these jobs. Analyzing complex logs of large systems is tedious. It is helpful to automatically analyze the logs in real-time and take reactive measures. In this paper, we present a unified job analysis and prediction system for supercomputer jobs. The users and administrators can monitor the current system state, analyze historical data and predict wait-times of future jobs. We evaluated our wait-time predictors on real job traces from 10 different systems. We observed 92.3% lower average prediction errors, as compared to existing methods. Ashish Pal, Preeti Malakar |
CLUSTER | 2 |
| 2021 | Parallel Program Scaling Analysis using Hardware CountersabstractWe present a lightweight library that automatically collects several hardware counters for MPI applications. We analyze the effect of strong and weak scaling on the counters. We first correlate the counter values obtained from each process count, and then cluster the counters to identify counters that are affected similarly due to scaling. We noted that the effect of last-level cache misses is more pronounced for some applications such as miniFE. Shobhit Jagga, Preeti Malakar |
HPDC | 2 |
| 2020 | MAP: A Visual Analytics System for Job Monitoring and AnalysisabstractHigh-performance computing systems are used for compute-intensive jobs by multiple users. They submit jobs to batch queues where the jobs are queued for an unknown amount of time until the required resources are available. A large amount of data is collected by the resource managers regarding the jobs (submit time, start time, end time, resource requirements, etc.). Analyzing this data may help identify causes of problems that may have occurred in the past and better optimize the system. Analyzing complex and huge logs may be cumbersome. We have developed a unified job monitoring, analysis, and prediction system using which users can monitor current state, analyze past job logs, and predict wait-times of future jobs. In this paper, we have focused on the job monitoring and analysis modules. Ashish Pal, Preeti Malakar |
CLUSTER | 2 |
| 2018 | Topology-aware space-shared co-analysis of large-scale molecular dynamics simulations
Preeti Malakar, Todd S. Munson, Christopher Knight 0001, Venkatram Vishwanath, Michael E. Papka |
SC | 1 |
| 2017 | Data movement optimizations for independent MPI I/O on the Blue Gene/Q
Preeti Malakar, Venkatram Vishwanath |
Parallel Comput. | 1 |
| 2016 | Optimal execution of co-analysis for large-scale molecular dynamics simulationsabstractThe analysis of scientific simulation data enables scientists to derive insights from their simulations. This analysis of the simulation output can be performed at the same execution site as the simulation using the same resources or can be done at a different site. The optimal output frequency is challenging to decide and is often chosen empirically. We propose a mathematical formulation for choosing the optimal frequency of data transfer for analysis and the feasibility of performing the analysis, under the given resource constraints such as network bandwidth, disk space, available memory, and computation time. We propose formulations for two cases of co-analysis - local and remote. We consider various analyses features such as computation time, input data and memory requirement, importance of the analysis and minimum frequency required for performing the analysis. We demonstrate the effectiveness of our approach using molecular dynamics applications on the Mira and Edison supercomputers. Preeti Malakar, Venkatram Vishwanath, Christopher Knight 0001, Todd S. Munson, Michael E. Papka |
SC | 1 |
| 2015 | Multipath Load Balancing for M × N Communication Patterns on the Blue Gene/Q Supercomputer Interconnection NetworkabstractAchievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually have a wide variation from application to application and even from phase to phase in an application, routing algorithms have a limited variation and usually are optimized for typical communication patterns. This results in high networking performance for favored communication patterns but low networking performance for others. In this paper we present approaches for improving networking performance by rebalancing load on physical links on the Blue Gene Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 30% higher throughput on experiment with data and patterns from a real application. The improvement can be up to several times higher throughput than default MPI_Alltoallv used in the Blue Gene Q supercomputer for certain communication patterns. Huy Bui, Robert L. Jacob, Preeti Malakar, Venkatram Vishwanath, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh |
CLUSTER | 3 |
| 2015 | Improving Communication Throughput by Multipath Load Balancing on Blue Gene/QabstractAchievable networking performance of applications in a supercomputer depends on the exact combination of the communication patterns of the applications and the routing algorithms used by the supercomputer. In order to achieve the highest networking performance for the applications, the routing algorithms need to be designed optimally for those communication patterns. However, while communication patterns usually vary from application to application and even from phase to phase in an application, routing algorithms have limited variation and usually are optimized for typical communication patterns. This results in high networking performance for some communication patterns. In this paper we present approaches for improving communication performance by using multiple paths and re-balancing load on physical links on the Blue Gene/Q supercomputer. We realize our approaches in a framework called OPTIQ and demonstrate the efficacy of our framework via a set of benchmarks. Our results show that we can achieve 43 -- 67% higher throughput on average from 91 experiments, and can achieve higher throughput than default MPI_Alltoallv used for certain communication patterns. Huy Bui, Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Eun-Sung Jung, Andrew E. Johnson 0001, Michael E. Papka, Jason Leigh |
HiPC | 2 |
| 2015 | Optimal scheduling of in-situ analysis for large-scale scientific simulationsabstractToday's leadership computing facilities have enabled the execution of transformative simulations at unprecedented scales. However, analyzing the huge amount of output from these simulations remains a challenge. Most analyses of this output is performed in post-processing mode at the end of the simulation. The time to read the output for the analysis can be significantly high due to poor I/O bandwidth, which increases the end-to-end simulation-analysis time. Simulation-time analysis can reduce this end-to-end time. In this work, we present the scheduling of in-situ analysis as a numerical optimization problem to maximize the number of online analyses subject to resource constraints such as I/O bandwidth, network bandwidth, rate of computation and available memory. We demonstrate the effectiveness of our approach through two application case studies on the IBM Blue Gene/Q system. Preeti Malakar, Venkatram Vishwanath, Todd S. Munson, Christopher Knight 0001, Mark Hereld, Sven Leyffer, Michael E. Papka |
SC | 1 |
| 2013 | A Diffusion-Based Processor Reallocation Strategy for Tracking Multiple Dynamically Varying Weather PhenomenaabstractMany meteorological phenomena occur at different locations simultaneously. These phenomena vary temporally and spatially. It is essential to track these multiple phenomena for accurate weather prediction. Efficient analysis require high-resolution simulations which can be conducted by introducing finer resolution nested simulations, nests at the locations of these phenomena. Simultaneous tracking of these multiple weather phenomena requires simultaneous execution of the nests on different subsets of the maximum number of processors for the main weather simulation. Dynamic variation in the number of these nests require efficient processor reallocation strategies. In this paper, we have developed strategies for efficient partitioning and repartitioning of the nests among the processors. As a case study, we consider an application of tracking multiple organized cloud clusters in tropical weather systems. We first present a parallel data analysis algorithm to detect such clouds. We have developed a tree-based hierarchical diffusion method which reallocates processors for the nests such that the redistribution cost is less. We achieve this by a novel tree reorganization approach. We show that our approach exhibits up to 25% lower redistribution cost and 53% lesser hop-bytes than the processor reallocation strategy that does not consider the existing processor allocation. Preeti Malakar, Vijay Natarajan, Sathish S. Vadhiyar, Ravi S. Nanjundiah |
ICPP | 1 |
| 2012 | Performance Evaluation and Optimization of Nested High Resolution Weather Simulations
Preeti Malakar, Vaibhav Saxena, Thomas George, Rashmi Mittal, Sameer Kumar 0001, Abdul Ghani Naim, Saiful Azmi bin Hj Husain |
Euro-Par | 1 |
| 2012 | A divide and conquer strategy for scaling weather simulations with multiple regions of interestabstractAccurate and timely prediction of weather phenomena, such as hurricanes and flash floods, require high-fidelity compute intensive simulations of multiple finer regions of interest within a coarse simulation domain. Current weather applications execute these nested simulations sequentially using all the available processors, which is sub-optimal due to their sub-linear scalability. In this work, we present a strategy for parallel execution of multiple nested domain simulations based on partitioning the 2-D processor grid into disjoint rectangular regions associated with each domain. We propose a novel combination of performance prediction, processor allocation methods and topology-aware mapping of the regions on torus interconnects. Experiments on IBM Blue Gene systems using WRF show that the proposed strategies result in performance improvement of up to 33% with topology-oblivious mapping and up to additional 7% with topology-aware mapping over the default sequential strategy. Preeti Malakar, Thomas George, Sameer Kumar 0001, Rashmi Mittal, Vijay Natarajan, Yogish Sabharwal, Vaibhav Saxena, Sathish S. Vadhiyar |
SC | 1 |
| 2010 | An Adaptive Framework for Simulation and Online Remote Visualization of Critical Climate Applications in Resource-constrained EnvironmentsabstractCritical climate applications like cyclone tracking and earthquake modeling require high-performance simulations and online visualization simultaneously performed with the simulations for timely analysis. Remote visualization of critical climate events enables joint analysis by geographically distributed climate science community. However, resource constraints including limited storage and slow networks can limit the effectiveness of such online visualization. In this work, we have developed an adaptive framework that simultaneously performs numerical simulations and online remote visualization of critical climate applications in resource-constrained environments. Our framework considers both application and resource dynamics to adapt various application and resource parameters including simulation resolutions, resource configurations and amount of data for visualization. We have developed two algorithms for processor allocation for simulations and the frequency of data for visualization. We show that our optimization method is able to provide about 30% higher simulation rate and consumes about 25-50% lesser storage space than the greedy approach. Preeti Malakar, Vijay Natarajan, Sathish S. Vadhiyar |
SC | 1 |