EDBT 2026 Demo / reviewers in the wild / expert
Abhishek Modi
dblp:188/9976
· DBLP profile ↗
4ranked-venue papers
1as first author
2since 2021 · last 2021
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 74% Indexing and storage engines · 22% Database system architecture and tuning · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 62% Performance modeling and evaluation · 38% | |
| Network and information security
1 paper |
Cryptographic protocols and secure computation · 50% Privacy and data protection · 50% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query optimization
distributed query optimization |
0.5 | 1 | 2021 | New Query Optimization Techniques in the Spark Engine of Azure Synapse · Proc. VLDB Endow. 2021 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.5 | 1 | 2021 | KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021 |
Performance modeling and evaluation
workload characterization |
0.5 | 1 | 2021 | KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021 |
Indexing and storage engines
i/o optimization |
0.4 | 1 | 2020 | Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data Queries · OSDI 2020 |
Privacy and data protection › privacy-preserving computation › encrypted data processing
encrypted data analytics |
0.2 | 1 | 2016 | Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016 |
Cryptographic protocols and secure computation
secure computation on encrypted data |
0.2 | 1 | 2016 | Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016 |
Cloud and datacenter computing
datacenter infrastructure |
0.1 | 1 | 2021 | KEA: Tuning an Exabyte-Scale Data Infrastructure · SIGMOD Conference 2021 |
Database system architecture and tuning › database security
encrypted data management |
0.1 | 1 | 2016 | Big Data Analytics over Encrypted Datasets with Seabed · OSDI 2016 |
Methods — techniques the papers use, named apart from their topics
peephole optimization · 1.0observational tuning · 0.5machine learning · 0.5flighting · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | KEA: Tuning an Exabyte-Scale Data InfrastructureabstractMicrosoft's internal big-data infrastructure is one of the largest in the world---with over 300k machines running billions of tasks from over 0.6M daily jobs. Operating this infrastructure is a costly and complex endeavor, and efficiency is paramount. In fact, for over 15 years, a dedicated engineering team has tuned almost every aspect of this infrastructure, achieving state-of-the-art efficiency (>60% average CPU utilization across all clusters). Despite rich telemetry and strong expertise, faced with evolving hardware/software/workloads this manual tuning approach had reached its limit---we had plateaued. In this paper, we present KEA, a multi-year effort to automate our tuning processes to be fully data/model-driven. KEA leverages a mix of domain knowledge and principled data science to capture the essence of our cluster dynamic behavior in a set of machine learning (ML) models based on collected system data. These models power automated optimization procedures for parameter tuning, and inform our leadership in critical decisions around engineering and capacity management (such as hardware and data center design, software investments, etc.). We combine "observational'' tuning (i.e., using models to predict system behavior without direct experimentation) with judicious use of "flighting'' (i.e., conservative testing in production). This allows us to support a broad range of applications that we discuss in this paper. KEA continuously tunes our cluster configurations and is on track to save Microsoft tens of millions of dollars per year. At the best of our knowledge, this paper is the first to discuss research challenges and practical learnings that emerge when tuning an exabyte-scale data infrastructure. Subru Krishnan, Konstantinos Karanasos, Isha Tarte, Conor Power, Abhishek Modi, Deli Zhang, Kartheek Muthyala, Nick Jurgens, Sarvesh Sakalanaga, Sudhir Darbha, Minu Iyer, Ankita Agarwal, Carlo Curino |
SIGMOD Conference | 6 |
| 2021 | New Query Optimization Techniques in the Spark Engine of Azure SynapseabstractThe cost of big-data query execution is dominated by stateful operators. These include sort and hash-aggregate that typically materialize intermediate data in memory, and exchange that materializes data to disk and transfers data over the network. In this paper we focus on several query optimization techniques that reduce the cost of these operators. First, we introduce a novel exchange placement algorithm that improves the state-of-the-art and significantly reduces the amount of data exchanged. The algorithm simultaneously minimizes the number of exchanges required and maximizes computation reuse via multi-consumer exchanges. Second, we introduce three partial push-down optimizations that push down partial computation derived from existing operators ( group-bys , intersections and joins ) below these stateful operators. While these optimizations are generically applicable we find that two of these optimizations ( partial aggregate and partial semi-join push-down ) are only beneficial in the scale-out setting where exchanges are a bottleneck. We propose novel extensions to existing literature to perform more aggressive partial push-downs than the state-of-the-art and also specialize them to the big-data setting. Finally we propose peephole optimizations that specialize the implementation of stateful operators to their input parameters. All our optimizations are implemented in the spark engine that powers azure synapse. We evaluate their impact on TPCDS and demonstrate that they make our engine 1.8X faster than Apache Spark 3.0.1. Abhishek Modi, Kaushik Rajan, Srinivas Thimmaiah, Prakhar Jain, Swinky Mann, Ayushi Agarwal, Ajith Shetty, Shahid K. I, Ashit Gosalia, Partho Sarthi |
Proc. VLDB Endow. | 1 |
| 2020 | Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data Queries
Partho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi, Prakhar Jain, Mo Liu 0003, Ashit Gosalia, Saurabh Kalikar |
OSDI | 4 |
| 2016 | Big Data Analytics over Encrypted Datasets with Seabed
Antonis Papadimitriou, Ranjita Bhagwan, Nishanth Chandran, Ramachandran Ramjee, Andreas Haeberlen, Harmeet Singh, Abhishek Modi, Saikrishna Badrinarayanan |
OSDI | 7 |