Parth Malani

dblp:80/1025 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0001-0589-5048ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Revisiting Reliability in Large-Scale Machine Learning Research Clusters
abstract
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to grow. Despite decades of research on infrastructure failures, the impact of job failures across different scales remains unclear. This paper presents a view of managing two large, multi-tenant ML clusters, providing quantitative analysis, operational experience, and our own perspective in understanding and addressing reliability concerns at scale. Our analysis reveals that while large jobs are most vulnerable to failures, smaller jobs make up the majority of jobs in the clusters and should be incorporated into optimization objectives. We identify key workload properties, compare them across clusters, and demonstrate essential reliability requirements for pushing the boundaries of ML training at scale.We hereby introduce a taxonomy of failures and key reliability metrics, analyze 11 months of data from two state-of-the-art ML environments with 4 million jobs and over 150 million A100 GPU hours. Building on our data, we fit a failure model to project Mean Time to Failure for various GPU scales. We further propose a method to estimate a related metric, Effective Training Time Ratio, as a function of job parameters, and we use this model to gauge the efficacy of potential software mitigations at scale. Our work provides valuable insights and future research directions for improving the reliability of AI supercomputer clusters, emphasizing the need for flexible, workload-agnostic, and reliability-aware infrastructure, system software, and algorithms.
Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zach DeVito, Shubho Sengupta, Kalyan Saladi, Carole-Jean Wu
HPCA5
2024 Expanding Datacenter Capacity with DVFS Boosting: A safe and scalable deployment experience
abstract
COVID-19 pandemic created unexpected demand for our physical infrastructure. We increased our computing supply by growing our infrastructure footprint as well as expanded existing capacity by using various techniques among those DVFS boosting. This paper describes our experience in deploying DVFS boosting to expand capacity.
Leonardo Piga, Iyswarya Narayanan, Aditya Sundarrajan, Matt Skach, Qingyuan Deng, Biswadip Maity, Manoj Chakkaravarthy, Alison Huang, Abhishek Dhanotia, Parth Malani
ASPLOS (1)10
2024 Dynamic Idle Resource Leasing To Safely Oversubscribe Capacity At Meta
abstract
Meta maintains additional capacity within its infrastructure to ensure high availability for business workloads, accommodating user growth, temporal traffic variations, and unforeseen regional failures. However, this strategic choice inherently leads to underutilization of resources. We employ oversubscription as an effective strategy to mitigate infrastructure underutilization.
Iyswarya Narayanan, Shivam Handa, Sayak Chakraborti, Pankit Thapar, Baohua Shan, Ariel Rao, Yuanlai Liu, Yuqing Wu, Qingyi Gao, Chris Chao-Chun Cheng, Sihan You, Louis Huang, Kenny Yu, Tengfei Mu, Parth Malani, Trey Lu, Peter Zhang
SoCC19
2023 Tutorial: MARS: A Framework for Runtime Monitoring, Modeling, and Management of Realtime Systems
Bryan Donyanavard, Nikil Dutt, Biswadip Maity, Parth Malani, Tiago Rogério Mück
CODES+ISSS4
2014 Accelerating pattern matching in neuromorphic text recognition system using Intel Xeon Phi coprocessor
abstract
Neuromorphic computing systems refer to the computing architecture inspired by the working mechanism of human brains. The rapidly reducing cost and increasing performance of state-of-the-art computing hardware allows large-scale implementation of machine intelligence models with neuromorphic architectures and opens the opportunity for new applications. One such computing hardware is Intel Xeon Phi coprocessor, which delivers over a TeraFLOP of computing power with 61 integrated processing cores. How to efficiently harness such computing power to achieve real time decision and cognition is one of the key design considerations. This paper presents an optimized implementation of Brain-State-in-a-Box (BSB) neural network model on the Xeon Phi coprocessor for pattern matching in the context of intelligent text recognition of noisy document images. From a scalability standpoint on a High Performance Computing (HPC) platform we show that efficient workload partitioning and resource management can double the performance of this many-core architecture for neuromorphic applications.
Khadeer Ahmed, Qinru Qiu, Parth Malani, Mangesh Tamhankar
IJCNN3
2010 Distributed task migration for thermal management in many-core systems
abstract
In the deep submicron era, thermal hot spots and large temperature gradients significantly impact system reliability, performance, cost and leakage power. As the system complexity increases, it is more and more difficult to perform thermal management in a centralized manner because of state explosion and the overhead of monitoring the entire chip. In this paper, we propose a framework for distributed thermal management for many-core systems where balanced thermal profile can be achieved by proactive task migration among neighboring cores. The framework has a low cost agent residing in each core that observes the local workload and temperature and communicates with its nearest neighbor for task migration/exchange. By choosing only those migration requests that will result balanced workload without generating thermal emergency, the proposed framework maintains workload balance across the system and avoids unnecessary migration. Experimental results show that, compared with existing proactive task migration technique, our approach generates less hotspots and smoother thermal gradient with less migration overhead and higher processing throughput.
Parth Malani, Qinru Qiu
DAC2
2008 Adaptive Scheduling and Voltage Scaling for Multiprocessor Real-time Applications with Non-deterministic Workload
abstract
The computational workload of some real-time applications varies significantly during runtime, which makes the task scheduling and power management a challenge. One of the major influences to the workload of an application is the selection of conditional branches which may activate or deactivate a large set of operations. Focusing on real-time applications with variable workload which is due to random branch selection, this paper presents a framework of task mapping, scheduling and dynamic voltage and frequency scaling (DVFS) for a multiprocessor system. The proposed framework maintains workload awareness using dynamic profiling of branch probability. The profiled information is utilized by the scheduling and DVFS algorithm that are adopted in this framework to generate statistically optimal solution.
Parth Malani, Prakash Mukre, Qinru Qiu, Qing Wu 0002
DATE1
2007 Resource-aware High Performance Scheduling for Embedded MPSoCs With the Application of MPEG Decoding
abstract
In this paper, we propose a scheduling algorithm to minimize the resource contentions and the processing latency for applications running on a multiprocessor system-on-chip (MPSoC) platform. The scheduling algorithm is applied on an MPSoC MPEG decoder to improve the system performance. Application specific task partition and mapping techniques are further investigated. The experimental results show an average improvement of 17% in total latency when comparing to the ad-hoc scheduled method.
Parth Malani, Qinru Qiu
ICME1
2007 Profile-Based Low Power Scheduling for Conditional Task Graph: A Communication Aware Approach
abstract
This work focuses on power optimization of realtime applications with conditional execution running on a dynamic voltage scaling (DVS) enabled multiprocessor system. A novel algorithm is proposed that performs simultaneous task mapping and ordering followed by task stretching of a conditional task graph (CTG). The algorithm minimizes the mathematical expectation of energy dissipation of non-deterministic applications with random branch selection by utilizing the task execution profile. Compared with existing scheduling algorithm, the experimental results show that our algorithm has 32% energy reduction in average.
Parth Malani, Prakash Mukre, Qinru Qiu
ISCAS1
2007 Power optimization for conditional task graphs in DVS enabled multiprocessor systems
abstract
In this paper, we focus on power optimization of real-time applications with conditional execution running on a dynamic voltage scaling (DVS) enabled multiprocessor system. The targeted system consists of heterogeneous processing elements with non-negligible inter-processor communication delay and energy. Given a conditional task graph (CTG), we have developed novel online and offline algorithms that perform simultaneous task mapping and ordering followed by task stretching. Both algorithms minimize the mathematical expectation of energy dissipation of non-deterministic applications by considering the probabilistic distribution of branch selection. Compared with existing CTG scheduling algorithms, our online and offline scheduling algorithms reduce energy by 28% and 39% in average, respectively.
Parth Malani, Prakash Mukre, Qinru Qiu
VLSI-SoC1
2006 Workload prediction and dynamic voltage scaling for MPEG decoding
abstract
In this paper we present three efficient DVS techniques for an MPEG decoder. Their energy reduction is comparable to that of the optimal solution. A workload prediction model is also developed based on the block level statistics of each MPEG frame. Compared with previous works, the new model exhibits a remarkable improvement in accuracy of the prediction. The experimental results show that, with the new prediction model, the presented DVS techniques achieve more energy reduction than previous works while delivering the same Quality of Service (QoS)
Parth Malani, Qinru Qiu, Qing Wu 0002
ASP-DAC2