VLDB 2026 Research / reviewers in the wild / expert
Edward Gan
dblp:95/11514
· DBLP profile ↗
11ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-3237-6657ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Query processing and optimization · 38% Data mining · 29% Data stream processing · 22% | |
| Theoretical computer science
1 paper |
Computational geometry · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
streaming analytics |
0.9 | 3 | 2018 | MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018 Demonstration: MacroBase, A Fast Data Analysis Engine · SIGMOD Conference 2017 MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017 |
Data integration and cleaning
data explanation |
0.8 | 2 | 2021 | DIFF: a relational interface for large-scale data explanation · VLDB J. 2021 DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018 |
Query processing and optimization
aggregation |
0.8 | 2 | 2020 | CoopStore: Optimizing Precomputed Summaries for Aggregation · Proc. VLDB Endow. 2020 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018 |
Data mining
anomaly detection |
0.6 | 2 | 2018 | MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018 MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017 |
Data mining › anomaly detection
streaming anomaly detection |
0.6 | 2 | 2018 | MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018 MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017 |
Query processing and optimization
query optimization |
0.4 | 1 | 2020 | CoopStore: Optimizing Precomputed Summaries for Aggregation · Proc. VLDB Endow. 2020 |
Query processing and optimization › aggregation
aggregate functions |
0.3 | 1 | 2018 | DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018 |
Data stream processing
quantile estimation |
0.3 | 1 | 2018 | Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018 |
Data stream processing › quantile estimation
quantile sketch |
0.3 | 1 | 2018 | Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018 |
Query processing and optimization
approximate query processing |
0.3 | 1 | 2017 | MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017 |
Data mining › clustering
density-based clustering |
0.3 | 1 | 2017 | Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017 |
Data mining › density estimation
kernel density estimation |
0.3 | 1 | 2017 | Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017 |
Computational geometry › spatial data structures
spatial indexing |
0.3 | 1 | 2017 | Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017 |
Systems and software security › isolation
software fault isolation |
0.1 | 1 | 2012 | RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012 |
Program verification › code-level verification
machine code verification |
0.1 | 1 | 2012 | RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012 |
Systems and software security › trusted computing
trusted computing base |
0.0 | 1 | 2012 | RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012 |
Methods — techniques the papers use, named apart from their topics
reservoir sampling · 0.6heavy-hitters sketch · 0.6classification · 0.6kernel density estimation · 0.6relational operators · 0.3method of moments · 0.3maximum entropy principle · 0.3feature selection · 0.3SQL extension · 0.3formal model in coq · 0.3declarative description · 0.3threshold-based pruning · 0.3streaming operators · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | DIFF: a relational interface for large-scale data explanation
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Ananthanarayan, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia |
VLDB J. | 4 |
| 2020 | CoopStore: Optimizing Precomputed Summaries for Aggregation
Edward Gan, Peter Bailis, Moses Charikar |
Proc. VLDB Endow. | 1 |
| 2020 | Approximate Selection with Guarantees using Proxies
Daniel Kang 0001, Edward Gan, Peter Bailis, Tatsunori B. Hashimoto, Matei Zaharia |
Proc. VLDB Endow. | 2 |
| 2018 | DIFF: A Relational Interface for Large-Scale Data ExplanationabstractA range of explanation engines assist data analysts by performing feature selection over increasingly high-volume and high-dimensional data, grouping and highlighting commonalities among data points. While useful in diverse tasks such as user behavior analytics, operational event processing, and root cause analysis, today's explanation engines are designed as standalone data processing tools that do not interoperate with traditional, SQL-based analytics workflows; this limits the applicability and extensibility of these engines. In response, we propose the DIFF operator, a relational aggregation operator that unifies the core functionality of these engines with declarative relational query processing. We implement both single-node and distributed versions of the DIFF operator in MB SQL, an extension of MacroBase, and demonstrate how DIFF can provide the same semantics as existing explanation engines while capturing a broad set of production use cases in industry, including at Microsoft and Facebook. Additionally, we illustrate how this declarative approach to data explanation enables new logical and physical query optimizations. We evaluate these optimizations on several real-world production applications, and find that DIFF in MB SQL can outperform state-of-the-art engines by up to an order of magnitude. Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Anathanaraya, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia |
Proc. VLDB Endow. | 4 |
| 2018 | Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation QueriesabstractInteractive analytics increasingly involves querying for quantiles over sub-populations of high cardinality datasets. Data processing engines such as Druid and Spark use mergeable summaries to estimate quantiles, but summary merge times can be a bottleneck during aggregation. We show how a compact and efficiently mergeable quantile sketch can support aggregation workloads. This data structure, which we refer to as the moments sketch, operates with a small memory footprint (200 bytes) and computationally efficient (50ns) merges by tracking only a set of summary statistics, notably the sample moments. We demonstrate how we can efficiently estimate quantiles using the method of moments and the maximum entropy principle, and show how the use of a cascade further improves query time for threshold predicates. Empirical evaluation shows that the moments sketch can achieve less than 1 percent quantile error with 15× less overhead than comparable summaries, improving end query time in the MacroBase engine by up to 7× and the Druid engine by up to 60×. Edward Gan, Jialin Ding 0001, Kai Sheng Tai, Vatsal Sharan, Peter Bailis |
Proc. VLDB Endow. | 1 |
| 2018 | MacroBase: Prioritizing Attention in Fast DataabstractAs data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation (i.e., feature selection) and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles. Firas Abuzaid, Peter Bailis, Jialin Ding 0001, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri |
ACM Trans. Database Syst. | 4 |
| 2017 | Prioritizing Attention in Analytic Monitoring
Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri |
CIDR | 2 |
| 2017 | MacroBase: Prioritizing Attention in Fast DataabstractAs data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles. Peter Bailis, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri |
SIGMOD Conference | 2 |
| 2017 | Demonstration: MacroBase, A Fast Data Analysis EngineabstractData volumes are rising at an increasing rate, stressing the limits of human attention. Current techniques for prioritizing user attention in this fast data are characterized by either cumbersome, ad-hoc analysis pipelines comprised of a diverse set of analytics tools, or brittle, static rule-based engines. To address this gap, we have developed MacroBase, a fast data analytics engine that acts as a search engine over fast data streams. MacroBase provides a set of highly-optimized, modular operators for streaming feature transformation, classification, and explanation. Users can leverage these optimized operators to construct efficient pipelines tailored for their use case. In this demonstration, SIGMOD attendees will have the opportunity to interactively answer and refine queries using MacroBase and discover the potential benefits of an advanced engine for prioritizing attention in high-volume, real-world data streams. Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri |
SIGMOD Conference | 2 |
| 2017 | Scalable Kernel Density Classification via Threshold-Based PruningabstractDensity estimation forms a critical component of many analytics tasks including outlier detection, visualization, and statistical testing. These tasks often seek to classify data into high and low-density regions of a probability distribution. Kernel Density Estimation (KDE) is a powerful technique for computing these densities, offering excellent statistical accuracy but quadratic total runtime. In this paper, we introduce a simple technique for improving the performance of using a KDE to classify points by their density (density classification). Our technique, thresholded kernel density classification (tKDC), applies threshold-based pruning to spatial index traversal to achieve asymptotic speedups over naïve KDE, while maintaining accuracy guarantees. Instead of exactly computing each point's exact density for use in classification, tKDC iteratively computes density bounds and short-circuits density computation as soon as bounds are either higher or lower than the target classification threshold. On a wide range of dataset sizes and dimensions, tKDC demonstrates empirical speedups of up to 1000x over alternatives. Edward Gan, Peter Bailis |
SIGMOD Conference | 1 |
| 2012 | RockSalt: better, faster, stronger SFI for the x86abstractSoftware-based fault isolation (SFI), as used in Google's Native Client (NaCl), relies upon a conceptually simple machine-code analysis to enforce a security policy. But for complicated architectures such as the x86, it is all too easy to get the details of the analysis wrong. We have built a new checker that is smaller, faster, and has a much reduced trusted computing base when compared to Google's original analysis. The key to our approach is automatically generating the bulk of the analysis from a declarative description which we relate to a formal model of a subset of the x86 instruction set architecture. The x86 model, developed in Coq, is of independent interest and should be usable for a wide range of machine-level verification tasks. J. Gregory Morrisett, Gang Tan, Joseph Tassarotti, Jean-Baptiste Tristan, Edward Gan |
PLDI | 5 |