Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Abhishek Modi

dblp:188/9976 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2021
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 74% Indexing and storage engines · 22% Database system architecture and tuning · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 62% Performance modeling and evaluation · 38%
Network and information security
1 paper
Cryptographic protocols and secure computation · 50% Privacy and data protection · 50%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query optimization
distributed query optimization
0.512021
New Query Optimization Techniques in the Spark Engine of Azure Synapse · Proc. VLDB Endow. 2021
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.512021
KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021
Performance modeling and evaluation
workload characterization
0.512021
KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021
Indexing and storage engines
i/o optimization
0.412020
Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data Queries · OSDI 2020
Privacy and data protection › privacy-preserving computation › encrypted data processing
encrypted data analytics
0.212016
Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016
Cryptographic protocols and secure computation
secure computation on encrypted data
0.212016
Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016
Cloud and datacenter computing
datacenter infrastructure
0.112021
KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021
Database system architecture and tuning › database security
encrypted data management
0.112016
Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016

Methods — techniques the papers use, named apart from their topics

peephole optimization · 1.0observational tuning · 0.5machine learning · 0.5flighting · 0.5
YearPublicationVenuePosition
2021 KEA: Tuning an Exabyte-Scale Data Infrastructure
abstract
Microsoft's internal big-data infrastructure is one of the largest in the world---with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infrastructure is a costly and complex endeavor, and efficiency is paramount. In fact, for over 15 years, a dedicated engineering team has tuned almost every aspect of this infrastructure, achieving state-of-the-art efficiency (>60% average CPU utilization across all clusters). Despite rich telemetry and strong expertise, faced with evolving hardware/software/workloads this manual tuning approach had reached its limit---we had plateaued. In this paper, we present KEA, a multi-year effort to automate our tuning processes to be fully data/model-driven. KEA leverages a mix of domain knowledge and principled data science to capture the essence of our cluster dynamic behavior in a set of machine learning (ML) models based on collected system data. These models power automated optimization procedures for parameter tuning, and inform our leadership in critical decisions around engineering and capacity management (such as hardware and data center design, software investments, etc.). We combine "observational'' tuning (i.e., using models to predict system behavior without direct experimentation) with judicious use of "flighting'' (i.e., conservative testing in production). This allows us to support a broad range of applications that we discuss in this paper. KEA continuously tunes our cluster configurations and is on track to save Microsoft tens of millions of dollars per year. At the best of our knowledge, this paper is the first to discuss research challenges and practical learnings that emerge when tuning an exabyte-scale data infrastructure.
Subru Krishnan, Konstantinos Karanasos, Isha Tarte, Conor Power, Abhishek Modi, Deli Zhang, Kartheek Muthyala, Nick Jurgens, Sarvesh Sakalanaga, Sudhir Darbha, Minu Iyer, Ankita Agarwal, Carlo Curino
SIGMOD Conference6
2021 New Query Optimization Techniques in the Spark Engine of Azure Synapse
abstract
The cost of big-data query execution is dominated by stateful operators. These include sort and hash-aggregate that typically materialize intermediate data in memory, and exchange that materializes data to disk and transfers data over the network. In this paper we focus on several query optimization techniques that reduce the cost of these operators. First, we introduce a novel exchange placement algorithm that improves the state-of-the-art and significantly reduces the amount of data exchanged. The algorithm simultaneously minimizes the number of exchanges required and maximizes computation reuse via multi-consumer exchanges. Second, we introduce three partial push-down optimizations that push down partial computation derived from existing operators ( group-bys , intersections and joins ) below these stateful operators. While these optimizations are generically applicable we find that two of these optimizations ( partial aggregate and partial semi-join push-down ) are only beneficial in the scale-out setting where exchanges are a bottleneck. We propose novel extensions to existing literature to perform more aggressive partial push-downs than the state-of-the-art and also specialize them to the big-data setting. Finally we propose peephole optimizations that specialize the implementation of stateful operators to their input parameters. All our optimizations are implemented in the spark engine that powers azure synapse. We evaluate their impact on TPCDS and demonstrate that they make our engine 1.8X faster than Apache Spark 3.0.1.
Abhishek Modi, Kaushik Rajan, Srinivas Thimmaiah, Prakhar Jain, Swinky Mann, Ayushi Agarwal, Ajith Shetty, Shahid K. I, Ashit Gosalia, Partho Sarthi
Proc. VLDB Endow.1
2020 Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data Queries
Partho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi, Prakhar Jain, Mo Liu 0003, Ashit Gosalia, Saurabh Kalikar
OSDI4
2016 Big Data Analytics over Encrypted Datasets with Seabed
Antonis Papadimitriou, Ranjita Bhagwan, Nishanth Chandran, Ramachandran Ramjee, Andreas Haeberlen, Harmeet Singh, Abhishek Modi, Saikrishna Badrinarayanan
OSDI7