Mayank Pundir

dblp:83/10462 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-3258-3357ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences
Neeraj Kumar 0004, Pol Mauri Ruiz, Igor Kabiljo, Mayank Pundir, Andrew Newell, Chunqiang Tang
OSDI5
2023 HHVM Performance Optimization for Large Scale Web Services
abstract
HHVM is commonly developed for large online web services, yet there remains much room for optimizing HHVM performance. This paper discusses challenges and techniques in optimizing HHVM performance for Meta's web service. We begin by evaluating the effectiveness of semantic request routing, a request routing method aimed at enhancing code cache performance in HHVM, and examine its implications for optimizing HHVM performance. Second, we characterize HHVM performance for a large-scale datacenter and identify the challenges brought by uncontrollable confounding factors. Finally, we present the performance management framework for autotuning HHVM performance at scale.
Alex Yang, Peinan Chen, Joey Pinto, Brian Karrer, Mayank Pundir, Maximilian Balandat, Arun Kejariwal, Benjamin C. Lee
ICPE7
2021 RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation
abstract
Capacity reservation is a common offering in public clouds and on-premise infrastructure. However, no prior work provides capacity reservation with SLO guarantees that takes into account random and correlated hardware failures, datacenter maintenance, and heterogeneous hardware. In this paper, we describe how Facebook's region-scale Resource Allowance System (RAS) addresses these issues and provides guaranteed capacity. RAS uses a capacity abstraction called reservation to represent a set of servers dynamically assigned to a logical cluster. We take a two-level approach to scale resource allocation to all datacenters in a region, where a mixed-integer-programming solver continuously optimizes server-to-reservation assignments off the critical path, and a traditional container allocator does real-time placement of containers on servers in a reservation. As a relatively new component of Facebook's 10-year old cluster manager Twine, RAS has been running in production for almost two years, continuously optimizing the allocation of millions of servers to thousands of reservations. We describe the design of RAS and share our experience of deploying it at scale.
Andrew Newell, Dimitrios Skarlatos 0002, Maxim Khutornenko, Mayank Pundir, Yuanlai Liu, Linh Le, Brendon Daugherty, Apurva Samudra, Prashasti Baid, James Kneeland, Igor Kabiljo, Dmitry Shchukin, Andre Rodrigues, Scott Michelson, Ben Christensen, Kaushik Veeraraghavan, Chunqiang Tang
SOSP6
2017 Social Hash Partitioner: A Scalable Distributed Hypergraph Partitioner
abstract
We design and implement a distributed algorithm for balanced k -way hypergraph partitioning that minimizes fanout, a fundamental hypergraph quantity also known as the communication volume and ( k - 1)-cut metric, by optimizing a novel objective called probabilistic fanout. This choice allows a simple local search heuristic to achieve comparable solution quality to the best existing hypergraph partitioners. Our algorithm is arbitrarily scalable due to a careful design that controls computational complexity, space complexity, and communication. In practice, we commonly process hypergraphs with billions of vertices and hyperedges in a few hours. We explain how the algorithm's scalability, both in terms of hypergraph size and bucket count, is limited only by the number of machines available. We perform an extensive comparison to existing distributed hypergraph partitioners and find that our approach is able to optimize hypergraphs roughly 100 times bigger on the same set of machines. We call the resulting tool Social Hash Partitioner , and accompanying this paper, we open-source the most scalable version based on recursive bisection.
Igor Kabiljo, Brian Karrer, Mayank Pundir, Sergey Pupyrev, Alon Shalita, Yaroslav Akhremtsev, Alessandro Presta
Proc. VLDB Endow.3
2017 Pandas: Robust Locality-Aware Scheduling With Stochastic Delay Optimality
abstract
Data locality is a fundamental problem to data-parallel applications where data-processing tasks consume different amounts of time and resources at different locations. The problem is especially prominent under stressed conditions such as hot spots. While replication based on data popularity relieves hot spots due to contention for a single file, hot spots caused by skewed node popularity, due to contention for files co-located with each other, are more complex, unpredictable, hence more difficult to deal with. We propose Pandas, a light-weight acceleration engine for data-processing tasks that is robust to changes in load and skewness in node popularity. Pandas is a stochastic delay-optimal algorithm. Trace-driven experiments on Hadoop show that Pandas accelerates the data-processing phase of jobs by 11 times with hot spots and 2.4 times without hot spots over existing schedulers. When the difference in processing times due to location is large, such as applicable to the case of memory-locality, the acceleration by Pandas is 22 times.
Qiaomin Xie, Mayank Pundir, Yi Lu 0001, Cristina L. Abad, Roy H. Campbell
IEEE/ACM Trans. Netw.2
2016 Supporting On-demand Elasticity in Distributed Graph Processing
abstract
While distributed graph processing engines have become popular for processing large graphs, these engines are typically configured with a static set of servers in the cluster. In other words, they lack the flexibility to scale-out or scale-in the number of servers, when requested to do so by the user. In this paper, we propose the first techniques to make distributed graph processing truly elastic. While supporting on-demand scale-out/in operations, we meet three goals: i) perform scale-out/in without interrupting the graph computation, ii) minimize the background network overhead involved in the scale-out/in, and iii) mitigate stragglers by maintaining load balance across servers. We present and analyze two techniques called Contiguous Vertex Repartitioning (CVR) and Ring-based Vertex Repartitioning (RVR) to address these goals. We implement our techniques in the LFGraph distributed graph processing system, and incorporate several systems optimizations. Experiments performed with multiple graph benchmark applications on a real graph indicate that our techniques perform within 9% and 21% of the optimum for scale-out and scale-in operations, respectively.
Mayank Pundir, Luke M. Leslie, Indranil Gupta, Roy H. Campbell
IC2E1
2015 Zorro: zero-cost reactive failure recovery in distributed graph processing
abstract
Distributed graph processing systems largely rely on proactive techniques for failure recovery. Unfortunately, these approaches (such as checkpointing) entail a significant overhead. In this paper, we argue that distributed graph processing systems should instead use a reactive approach to failure recovery. The reactive approach trades off completeness of the result (generating a slightly inaccurate result) while reducing the overhead during failure-free execution to zero. We build a system called Zorro that imbues this reactive approach, and integrate Zorro into two graph processing systems -- PowerGraph and LFGraph. When a failure occurs, Zorro opportunistically exploits vertex replication inherent in today's graph processing systems to quickly rebuild the state of failed servers. Experiments using real-world graphs demonstrate that Zorro is able to recover over 99% of the graph state when 6--12% of the servers fail, and between 87--95% when half the cluster fails. Furthermore, using various graph processing algorithms, Zorro incurs little to no accuracy loss in all experimental failure scenarios, and achieves a worst-case accuracy of 97%.
Mayank Pundir, Luke M. Leslie, Indranil Gupta, Roy H. Campbell
SoCC1
2012 Work in progress: On entrance test criteria for CS and IT UG programs
abstract
Indian universities broadly follow an entrance process for undergraduate engineering process in which physics, chemistry and mathematics based questions are asked. Either the individual scores or some combination of these scores is used to set the admission criterion for different engineering streams. However, it has not been analyzed whether this is a good predictor of undergraduate performance or not, especially in Indian context. The broad goal of this research is to understand the relationship between entrance criterion and undergraduate Computer Science/Information Technology performance. The analysis includes data collected from the students who appeared in the entrance exams and took admission in IIIT-Delhi in between 2008 to 2010. The scores obtained in entrance exams and cumulative grade point average earned during the first year of their studies are analyzed. Preliminary analysis suggests that for computer science undergraduate studies, mathematics or numerical understanding is an important aspect. Further, logic and aptitude based tests are better correlated in context of computer science education.
Richa Singh 0001, Mayank Pundir
FIE2