EDBT 2026 Demo / reviewers in the wild / expert
Andrew Fogarty
dblp:273/7169
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0000-3248-655XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 73% High-performance computing · 13% Distributed systems · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Indexing and storage engines · 32% Query processing and optimization · 32% Data stream processing · 28% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.8 | 1 | 2024 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing › resource allocation › dynamic resource allocation
proactive resource allocation |
0.8 | 1 | 2024 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing
resource provisioning |
0.8 | 1 | 2024 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service · Proc. VLDB Endow. 2024 |
Query processing and optimization › query execution
index-based query processing |
0.5 | 1 | 2021 | Hyperspace: The Indexing Subsystem of Azure Synapse · Proc. VLDB Endow. 2021 |
Indexing and storage engines
index management |
0.5 | 1 | 2021 | Hyperspace: The Indexing Subsystem of Azure Synapse · Proc. VLDB Endow. 2021 |
Distributed systems › distributed data processing
distributed indexing |
0.4 | 1 | 2020 | Helios: Hyperscale Indexing for the Cloud & Edge · Proc. VLDB Endow. 2020 |
High-performance computing › data-intensive computing
large-scale data processing |
0.4 | 1 | 2020 | Helios: Hyperscale Indexing for the Cloud & Edge · Proc. VLDB Endow. 2020 |
Methods — techniques the papers use, named apart from their topics
distributed systems design · 0.9hyperparameter auto-tuning · 0.8hybrid machine learning model · 0.8concurrency control · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud ServiceabstractThe proliferation of big data and analytic workloads has driven the need for cloud compute and cluster-based job processing. With Apache Spark, users can process terabytes of data at ease with hundreds of parallel executors. Providing low latency access to Spark clusters and sessions is a challenging problem due to the large overheads of cluster creation and session startup. In this paper, we introduce Intelligent Pooling, a system for proactively provisioning compute resources to combat the aforementioned overheads. Our system (1) predicts usage patterns using an innovative hybrid Machine Learning (ML) model with low latency and high accuracy; and (2) optimizes the pool size dynamically to meet customer demand while reducing extraneous COGS. The proposed system auto-tunes its hyper-parameters to balance between performance and operational cost with minimal to no engineering input. Evaluated using large-scale production data, Intelligent Pooling achieves up to 43% reduction in cluster idle time compared to static pooling when targeting 99% pool hit rate. Currently deployed in production, Intelligent Pooling is on track to save tens of million dollars in COGS per year as compared to traditional pre-provisioned pools. Deepak Ravikumar, Alex Yeo, Aditya Lakra, Harsha Nagulapalli, Santhosh Ravindran, Steve Suh, Niharika Dutta, Andrew Fogarty, Yoonjae Park, Sumeet Khushalani, Arijit Tarafdar, Kunal Parekh, Subru Krishnan |
Proc. VLDB Endow. | 9 |
| 2021 | Hyperspace: The Indexing Subsystem of Azure SynapseabstractMicrosoft recently introduced Azure Synapse Analytics, which offers an integrated experience across data ingestion, storage, and querying in Apache Spark and T-SQL over data in the lake, including files and warehouse tables. In this paper, we present our experiences with designing and implementing Hyperspace, the indexing subsystem underlying Synapse. Hyperspace enables users to build multiple types of secondary indexes on their data, maintain them through a multi-user concurrency model, and leverage them automatically---without any change to their application code---for query/workload acceleration. Many requirements of Hyperspace are based on feedback from several enterprise customers. We present the details of Hyperspace's underlying design, the user-facing APIs, its concurrency control protocol for index access, its index-aware query processing techniques, and its maintenance mechanisms for handling index updates. Evaluations over standard industry benchmarks and real customer workloads show that Hyperspace can accelerate query execution by up to 10x and in certain real-world workloads, even up to two orders of magnitude. Rahul Potharaju, Terry Kim, Eunjin Song, Wentao Wu 0001, Lev Novik, Apoorve Dave, Pouria Pirzadeh, Andrew Fogarty, Gurleen Dhody, Jiying Li, Vidip Acharya, Sinduja Ramanujam, Nicolas Bruno, César A. Galindo-Legaria, Vivek R. Narasayya, Surajit Chaudhuri, Anil K. Nori, Tomas Talius, Raghu Ramakrishnan 0001 |
Proc. VLDB Endow. | 8 |
| 2020 | Helios: Hyperscale Indexing for the Cloud & EdgeabstractHelios is a distributed, highly-scalable system used at Microsoft for flexible ingestion, indexing, and aggregation of large streams of real-time data that is designed to plug into relational engines. The system collects close to a quadrillion events indexing approximately 16 trillion search keys per day from hundreds of thousands of machines across tens of data centers around the world. Helios use cases within Microsoft include debugging/diagnostics in both public and government clouds, workload characterization, cluster health monitoring, deriving business insights and performing impact analysis of incidents in other large-scale systems such as Azure Data Lake and Cosmos. Helios also serves as a reference blueprint for other large-scale systems within Microsoft. We present the simple data model behind Helios, which offers great flexibility and control over costs, and enables the system to asynchronously index massive streams of data. We also present our experiences in building and operating Helios over the last five years at Microsoft. Rahul Potharaju, Terry Kim, Wentao Wu 0001, Vidip Acharya, Steve Suh, Andrew Fogarty, Apoorve Dave, Sinduja Ramanujam, Tomas Talius, Lev Novik, Raghu Ramakrishnan 0001 |
Proc. VLDB Endow. | 6 |