Edward Gan

dblp:95/11514 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-3237-6657ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 38% Data mining · 29% Data stream processing · 22%
Theoretical computer science
1 paper
Computational geometry · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
streaming analytics
0.932018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
Demonstration: MacroBase, A Fast Data Analysis Engine · SIGMOD Conference 2017
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Data integration and cleaning
data explanation
0.822021
DIFF: a relational interface for large-scale data explanation · VLDB J. 2021
DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018
Query processing and optimization
aggregation
0.822020
CoopStore: Optimizing Precomputed Summaries for Aggregation · Proc. VLDB Endow. 2020
Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018
Data mining
anomaly detection
0.622018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Data mining › anomaly detection
streaming anomaly detection
0.622018
MacroBase: Prioritizing Attention in Fast Data · ACM Trans. Database Syst. 2018
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Query processing and optimization
query optimization
0.412020
CoopStore: Optimizing Precomputed Summaries for Aggregation · Proc. VLDB Endow. 2020
Query processing and optimization › aggregation
aggregate functions
0.312018
DIFF: A Relational Interface for Large-Scale Data Explanation · Proc. VLDB Endow. 2018
Data stream processing
quantile estimation
0.312018
Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018
Data stream processing › quantile estimation
quantile sketch
0.312018
Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries · Proc. VLDB Endow. 2018
Query processing and optimization
approximate query processing
0.312017
MacroBase: Prioritizing Attention in Fast Data · SIGMOD Conference 2017
Data mining › clustering
density-based clustering
0.312017
Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017
Data mining › density estimation
kernel density estimation
0.312017
Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017
Computational geometry › spatial data structures
spatial indexing
0.312017
Scalable Kernel Density Classification via Threshold-Based Pruning · SIGMOD Conference 2017
Systems and software security › isolation
software fault isolation
0.112012
RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012
Program verification › code-level verification
machine code verification
0.112012
RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012
Systems and software security › trusted computing
trusted computing base
0.012012
RockSalt: better, faster, stronger SFI for the x86 · PLDI 2012

Methods — techniques the papers use, named apart from their topics

reservoir sampling · 0.6heavy-hitters sketch · 0.6classification · 0.6kernel density estimation · 0.6relational operators · 0.3method of moments · 0.3maximum entropy principle · 0.3feature selection · 0.3SQL extension · 0.3formal model in coq · 0.3declarative description · 0.3threshold-based pruning · 0.3streaming operators · 0.3
YearPublicationVenuePosition
2021 DIFF: a relational interface for large-scale data explanation
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Ananthanarayan, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
VLDB J.4
2020 CoopStore: Optimizing Precomputed Summaries for Aggregation
Edward Gan, Peter Bailis, Moses Charikar
Proc. VLDB Endow.1
2020 Approximate Selection with Guarantees using Proxies
Daniel Kang 0001, Edward Gan, Peter Bailis, Tatsunori B. Hashimoto, Matei Zaharia
Proc. VLDB Endow.2
2018 DIFF: A Relational Interface for Large-Scale Data Explanation
abstract
A range of explanation engines assist data analysts by performing feature selection over increasingly high-volume and high-dimensional data, grouping and highlighting commonalities among data points. While useful in diverse tasks such as user behavior analytics, operational event processing, and root cause analysis, today's explanation engines are designed as standalone data processing tools that do not interoperate with traditional, SQL-based analytics workflows; this limits the applicability and extensibility of these engines. In response, we propose the DIFF operator, a relational aggregation operator that unifies the core functionality of these engines with declarative relational query processing. We implement both single-node and distributed versions of the DIFF operator in MB SQL, an extension of MacroBase, and demonstrate how DIFF can provide the same semantics as existing explanation engines while capturing a broad set of production use cases in industry, including at Microsoft and Facebook. Additionally, we illustrate how this declarative approach to data explanation enables new logical and physical query optimizations. We evaluate these optimizations on several real-world production applications, and find that DIFF in MB SQL can outperform state-of-the-art engines by up to an order of magnitude.
Firas Abuzaid, Peter Kraft, Sahaana Suri, Edward Gan, Eric Xu, Atul Shenoy, Asvin Anathanaraya, John Sheu, Erik Meijer 0001, Xi Wu 0001, Jeffrey F. Naughton, Peter Bailis, Matei Zaharia
Proc. VLDB Endow.4
2018 Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries
abstract
Interactive analytics increasingly involves querying for quantiles over sub-populations of high cardinality datasets. Data processing engines such as Druid and Spark use mergeable summaries to estimate quantiles, but summary merge times can be a bottleneck during aggregation. We show how a compact and efficiently mergeable quantile sketch can support aggregation workloads. This data structure, which we refer to as the moments sketch, operates with a small memory footprint (200 bytes) and computationally efficient (50ns) merges by tracking only a set of summary statistics, notably the sample moments. We demonstrate how we can efficiently estimate quantiles using the method of moments and the maximum entropy principle, and show how the use of a cascade further improves query time for threshold predicates. Empirical evaluation shows that the moments sketch can achieve less than 1 percent quantile error with 15× less overhead than comparable summaries, improving end query time in the MacroBase engine by up to 7× and the Druid engine by up to 60×.
Edward Gan, Jialin Ding 0001, Kai Sheng Tai, Vatsal Sharan, Peter Bailis
Proc. VLDB Endow.1
2018 MacroBase: Prioritizing Attention in Fast Data
abstract
As data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation (i.e., feature selection) and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles.
Firas Abuzaid, Peter Bailis, Jialin Ding 0001, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri
ACM Trans. Database Syst.4
2017 Prioritizing Attention in Analytic Monitoring
Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri
CIDR2
2017 MacroBase: Prioritizing Attention in Fast Data
abstract
As data volumes continue to rise, manual inspection is becoming increasingly untenable. In response, we present MacroBase, a data analytics engine that prioritizes end-user attention in high-volume fast data streams. MacroBase enables efficient, accurate, and modular analyses that highlight and aggregate important and unusual behavior, acting as a search engine for fast data. MacroBase is able to deliver order-of-magnitude speedups over alternatives by optimizing the combination of explanation and classification tasks and by leveraging a new reservoir sampler and heavy-hitters sketch specialized for fast data streams. As a result, MacroBase delivers accurate results at speeds of up to 2M events per second per query on a single core. The system has delivered meaningful results in production, including at a telematics company monitoring hundreds of thousands of vehicles.
Peter Bailis, Edward Gan, Samuel Madden 0001, Deepak Narayanan, Kexin Rong 0001, Sahaana Suri
SIGMOD Conference2
2017 Demonstration: MacroBase, A Fast Data Analysis Engine
abstract
Data volumes are rising at an increasing rate, stressing the limits of human attention. Current techniques for prioritizing user attention in this fast data are characterized by either cumbersome, ad-hoc analysis pipelines comprised of a diverse set of analytics tools, or brittle, static rule-based engines. To address this gap, we have developed MacroBase, a fast data analytics engine that acts as a search engine over fast data streams. MacroBase provides a set of highly-optimized, modular operators for streaming feature transformation, classification, and explanation. Users can leverage these optimized operators to construct efficient pipelines tailored for their use case. In this demonstration, SIGMOD attendees will have the opportunity to interactively answer and refine queries using MacroBase and discover the potential benefits of an advanced engine for prioritizing attention in high-volume, real-world data streams.
Peter Bailis, Edward Gan, Kexin Rong 0001, Sahaana Suri
SIGMOD Conference2
2017 Scalable Kernel Density Classification via Threshold-Based Pruning
abstract
Density estimation forms a critical component of many analytics tasks including outlier detection, visualization, and statistical testing. These tasks often seek to classify data into high and low-density regions of a probability distribution. Kernel Density Estimation (KDE) is a powerful technique for computing these densities, offering excellent statistical accuracy but quadratic total runtime. In this paper, we introduce a simple technique for improving the performance of using a KDE to classify points by their density (density classification). Our technique, thresholded kernel density classification (tKDC), applies threshold-based pruning to spatial index traversal to achieve asymptotic speedups over naïve KDE, while maintaining accuracy guarantees. Instead of exactly computing each point's exact density for use in classification, tKDC iteratively computes density bounds and short-circuits density computation as soon as bounds are either higher or lower than the target classification threshold. On a wide range of dataset sizes and dimensions, tKDC demonstrates empirical speedups of up to 1000x over alternatives.
Edward Gan, Peter Bailis
SIGMOD Conference1
2012 RockSalt: better, faster, stronger SFI for the x86
abstract
Software-based fault isolation (SFI), as used in Google's Native Client (NaCl), relies upon a conceptually simple machine-code analysis to enforce a security policy. But for complicated architectures such as the x86, it is all too easy to get the details of the analysis wrong. We have built a new checker that is smaller, faster, and has a much reduced trusted computing base when compared to Google's original analysis. The key to our approach is automatically generating the bulk of the analysis from a declarative description which we relate to a formal model of a subset of the x86 instruction set architecture. The x86 model, developed in Coq, is of independent interest and should be usable for a wide range of machine-level verification tasks.
J. Gregory Morrisett, Gang Tan, Joseph Tassarotti, Jean-Baptiste Tristan, Edward Gan
PLDI5