Emanuel Zgraggen

dblp:153/7539 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Query processing and optimization · 43% Data mining · 26% Data integration and cleaning · 12%
Computer graphics and multimedia
8 papers
Visualization and visual analytics · 100%
Human-computer interaction and pervasive computing
5 papers
Usability and user experience research · 50% Interaction techniques and input · 26% Human-AI interaction · 15%
Artificial intelligence
2 papers
Deep learning architectures and training · 79% Learning theory · 21%

Topics — the 27 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
1.132021
Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021
How Progressive Visualizations Affect Exploratory Analysis · IEEE Trans. Vis. Comput. Graph. 2017
Revisiting Reuse for Approximate Query Processing · Proc. VLDB Endow. 2017
Query processing and optimization › approximate query processing
sampling-based approximate query processing
0.512021
Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021
Visualization and visual analytics › interactive data exploration
visual exploration
0.522017
Safe Visual Data Exploration · SIGMOD Conference 2017
PanoramicData: Data Analysis through Pen & Touch · IEEE Trans. Vis. Comput. Graph. 2014
Database system architecture and tuning
database benchmarking
0.412020
IDEBench: A Benchmark for Interactive Data Exploration · SIGMOD Conference 2020
Data mining › dimensionality reduction
feature selection
0.412020
ARDA: Automatic Relational Data Augmentation for Machine Learning · Proc. VLDB Endow. 2020
Query processing and optimization
interactive data exploration
0.412020
IDEBench: A Benchmark for Interactive Data Exploration · SIGMOD Conference 2020
Data mining
pattern mining
0.422018
ProSecCo: Progressive Sequence Mining with Convergence Guarantees · ICDM 2018
Safe Visual Data Exploration · SIGMOD Conference 2017
Data integration and cleaning
data preprocessing
0.412019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019
Machine learning and data management
machine learning pipeline
0.412019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019
Data integration and cleaning › table understanding › table annotation
semantic type detection
0.412019
Sherlock: A Deep Learning Approach to Semantic Data Type Detection · KDD 2019
Visualization and visual analytics › visualization design
automated visualization
0.412019
VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository · CHI 2019
Visualization and visual analytics
visualization design
0.412019
VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository · CHI 2019
Visualization and visual analytics › visualization literacy
visualization education
0.412019
VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository · CHI 2019
Data mining › pattern mining › sequential pattern mining
frequent sequence mining
0.312018
ProSecCo: Progressive Sequence Mining with Convergence Guarantees · ICDM 2018
Query processing and optimization
result reuse
0.312017
Revisiting Reuse for Approximate Query Processing · Proc. VLDB Endow. 2017
Data mining › statistical analysis
statistical inference
0.312017
Controlling False Discoveries During Interactive Data Exploration · SIGMOD Conference 2017
Visualization and visual analytics › visual analytics
exploratory data analysis
0.312017
How Progressive Visualizations Affect Exploratory Analysis · IEEE Trans. Vis. Comput. Graph. 2017
Visualization and visual analytics › interactive visualization
progressive visualization
0.312017
How Progressive Visualizations Affect Exploratory Analysis · IEEE Trans. Vis. Comput. Graph. 2017
Visualization and visual analytics › interaction techniques
visual query interface
0.212015
(s|qu)eries: Visual Regular Expressions for Querying and Exploring Event Sequences · CHI 2015
Interaction techniques and input › input modality › multimodal input
pen and touch input
0.212014
PanoramicData: Data Analysis through Pen & Touch · IEEE Trans. Vis. Comput. Graph. 2014
Data mining
predictive modeling
0.112020
ARDA: Automatic Relational Data Augmentation for Machine Learning · Proc. VLDB Endow. 2020
Human-AI interaction
interactive machine learning
0.112019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019
Machine learning › Learning theory › computational learning theory › VC theory
VC dimension
0.112018
ProSecCo: Progressive Sequence Mining with Convergence Guarantees · ICDM 2018
Visualization and visual analytics
interactive data exploration
0.112017
Controlling False Discoveries During Interactive Data Exploration · SIGMOD Conference 2017
Data models and query languages
regular expressions
0.112015
(s|qu)eries: Visual Regular Expressions for Querying and Exploring Event Sequences · CHI 2015
User interface design and tools › natural user interface
interactive whiteboard
0.112015
Vizdom: Interactive Analytics through Pen and Touch · Proc. VLDB Endow. 2015
Data models and query languages › query language
visual query language
0.112014
PanoramicData: Data Analysis through Pen & Touch · IEEE Trans. Vis. Comput. Graph. 2014

Methods — techniques the papers use, named apart from their topics

word embeddings · 0.8paragraph vector · 0.8model selection · 0.8hyperparameter tuning · 0.8deep neural network · 0.8controlled experiment · 0.7confirmatory analysis · 0.7think-aloud protocol · 0.6statistical hypothesis testing · 0.6multiple testing correction · 0.6interaction logs · 0.6sampling · 0.5progressive computation · 0.5incremental visualization · 0.4approximation techniques · 0.4user study · 0.4feature selection · 0.4data search and join · 0.4
YearPublicationVenuePosition
2021 Davos: A System for Interactive Data-Driven Decision Making
abstract
Recently, a new horizon in data analytics, prescriptive analytics, is becoming more and more important to make data-driven decisions. As opposed to the progress of democratizing data acquisition and access, making data-driven decisions remains a significant challenge for people without technical expertise. In this regard, existing tools for data analytics which were designed decades ago still present a high bar for domain experts, and removing this bar requires a fundamental rethinking of both interface and backend. At Einblick, an MIT/Brown spin-off based on the Northstar project, we have been building the next generation analytics tool in the last few years. To overcome the shortcomings of existing processing engines, we propose Davos , Einblick's novel backend. Davos combines aspects of progressive computation, approximate query processing and sampling, with a specific focus on supporting user-defined operations. Moreover, Davos optimizes multi-tenant scenarios to promote collaboration. Both empirical evaluation and user study verify that Davos can greatly empower data analytics for new needs.
Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Philipp Eichmann, Navid Karimeddiny, Charlie Meyer, Wesley Runnels, Tim Kraska
Proc. VLDB Endow.2
2020 IDEBench: A Benchmark for Interactive Data Exploration
abstract
In recent years, many query processing techniques have been developed to better support interactive data exploration (IDE) of large structured datasets. To evaluate and compare database engines in terms of how well they support such workloads, experimenters have mostly used self-designed evaluation procedures rather than established benchmarks. In this paper we argue that this is due to the fact that the workloads and metrics of popular analytical benchmarks such as TPC-H or TPC-DS were designed for traditional performance reporting scenarios, and do not capture distinctive IDE characteristics. Guided by the findings of several user studies we present a new benchmark called IDEBench, designed to evaluate database engines based on common IDE workflows and metrics that matter to the end-user. We demonstrate the applicability of IDEBench through a number of experiments with five different database engines, and present and discuss our findings.
Philipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim Kraska
SIGMOD Conference2
2020 ProSecCo: progressive sequence mining with convergence guarantees
Sacha Servan-Schreiber, Matteo Riondato, Emanuel Zgraggen
Knowl. Inf. Syst.3
2020 ARDA: Automatic Relational Data Augmentation for Machine Learning
abstract
Automatic machine learning (AML) is a family of techniques to automate the process of training predictive models, aiming to both improve performance and make machine learning more accessible. While many recent works have focused on aspects of the machine learning pipeline like model selection, hyperparameter tuning, and feature selection, relatively few works have focused on automatic data augmentation. Automatic data augmentation involves finding new features relevant to the user's predictive task with minimal "human-in-the-loop" involvement. We present ARDA, an end-to-end system that takes as input a dataset and a data repository, and outputs an augmented data set such that training a predictive model on this augmented dataset results in improved performance. Our system has two distinct components: (1) a framework to search and join data with the input data, based on various attributes of the input, and (2) an efficient feature selection algorithm that prunes out noisy or irrelevant features from the resulting join. We perform an extensive empirical evaluation of different system components and benchmark our feature selection algorithm on real-world datasets.
Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez, Tim Kraska, David R. Karger
Proc. VLDB Endow.3
2019 VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository
abstract
Researchers currently rely on ad hoc datasets to train automated visualization tools and evaluate the effectiveness of visualization designs. These exemplars often lack the characteristics of real-world datasets, and their one-off nature makes it difficult to compare different techniques. In this paper, we present VizNet: a large-scale corpus of over 31 million datasets compiled from open data repositories and online visualization galleries. On average, these datasets comprise 17 records over 3 dimensions and across the corpus, we find 51% of the dimensions record categorical data, 44% quantitative, and only 5% temporal. VizNet provides the necessary common baseline for comparing visualization design techniques, and developing benchmark models and algorithms for automating visual analysis. To demonstrate VizNet's utility as a platform for conducting online crowdsourced experiments at scale, we replicate a prior study assessing the influence of user task and data distribution on visual encoding effectiveness, and extend it by considering an additional task: outlier detection. To contend with running such studies at scale, we demonstrate how a metric of perceptual effectiveness can be learned from experimental results, and show its predictive power across test datasets.
Kevin Zeng Hu, Snehalkumar (Neil) S. Gaikwad, Madelon Hulsebos, Michiel A. Bakker, Emanuel Zgraggen, César A. Hidalgo 0001, Tim Kraska, Guoliang Li 0001, Arvind Satyanarayan, Çagatay Demiralp
CHI5
2019 Sherlock: A Deep Learning Approach to Semantic Data Type Detection
abstract
Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparation and analysis systems rely on dictionary lookups and regular expression matching to detect semantic types. However, these matching-based approaches often are not robust to dirty data and only detect a limited number of types. We introduce Sherlock, a multi-input deep neural network for detecting semantic types. We train Sherlock on $686,765$ data columns retrieved from the VizNet corpus by matching $78$ semantic types from DBpedia to column headers. We characterize each matched column with $1,588$ features describing the statistical properties, character distributions, word embeddings, and paragraph vectors of column values. Sherlock achieves a support-weighted F$_1$ score of $0.89$, exceeding that of machine learning baselines, dictionary and regular expression benchmarks, and the consensus of crowdsourced annotations.
Madelon Hulsebos, Kevin Zeng Hu, Michiel A. Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, César A. Hidalgo 0001
KDD4
2019 Democratizing Data Science through Interactive Curation of ML Pipelines
abstract
Statistical knowledge and domain expertise are key to extract actionable insights out of data, yet such skills rarely coexist together. In Machine Learning, high-quality results are only attainable via mindful data preprocessing, hyperparameter tuning and model selection. Domain experts are often overwhelmed by such complexity, de-facto inhibiting a wider adoption of ML techniques in other fields. Existing libraries that claim to solve this problem, still require well-trained practitioners. Those frameworks involve heavy data preparation steps and are often too slow for interactive feedback from the user, severely limiting the scope of such systems.
Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Ferdinand Kossmann, Philipp Eichmann, Yeounoh Chung, Carsten Binnig, Eli Upfal, Tim Kraska
SIGMOD Conference2
2018 Investigating the Effect of the Multiple Comparisons Problem in Visual Analysis
abstract
The goal of a visualization system is to facilitate dataset-driven insight discovery. But what if the insights are spurious? Features or patterns in visualizations can be perceived as relevant insights, even though they may arise from noise. We often compare visualizations to a mental image of what we are interested in: a particular trend, distribution or an unusual pattern. As more visualizations are examined and more comparisons are made, the probability of discovering spurious insights increases. This problem is well-known in Statistics as the multiple comparisons problem (MCP) but overlooked in visual analysis. We present a way to evaluate MCP in visualization tools by measuring the accuracy of user reported insights on synthetic datasets with known ground truth labels. In our experiment, over 60% of user insights were false. We show how a confirmatory analysis approach that accounts for all visual comparisons, insights and non-insights, can achieve similar results as one that requires a validation dataset.
Emanuel Zgraggen, Zheguang Zhao, Robert C. Zeleznik, Tim Kraska
CHI1
2018 ProSecCo: Progressive Sequence Mining with Convergence Guarantees
abstract
We present PROSECCO, an algorithm for the progressive mining of frequent sequences from large transactional datasets: it processes the dataset in blocks and outputs, after having analyzed each block, a high-quality approximation of the collection of frequent sequences. These intermediate results have strong probabilistic approximation guarantees and the final output is the exact collection of frequent sequences. Our correctness analysis uses the Vapnik-Chervonenkis (VC) dimension, a key concept from statistical learning theory. The results of our experimental evaluation of PROSECCO on real and artificial datasets show that it produces fast-converging high-quality results almost immediately. Its practical performance is even better than what is guaranteed by the theoretical analysis, and it can even be faster than existing state-of-the-art non-progressive algorithms.
Sacha Servan-Schreiber, Matteo Riondato, Emanuel Zgraggen
ICDM3
2017 Toward Sustainable Insights, or Why Polygamy is Bad for You
Carsten Binnig, Lorenzo De Stefani, Tim Kraska, Eli Upfal, Emanuel Zgraggen, Zheguang Zhao
CIDR5
2017 Controlling False Discoveries During Interactive Data Exploration
abstract
Recent tools for interactive data exploration significantly increase the chance that users make false discoveries. They allow users to (visually) examine many hypotheses and make inference with simple interactions, and thus incur the issue commonly known in statistics as the "multiple hypothesis testing error." In this work, we propose a solution to integrate the control of multiple hypothesis testing into interactive data exploration systems. A key insight is that existing methods for controlling the false discovery rate (such as FDR) are not directly applicable to interactive data exploration. We therefore discuss a set of new control procedures that are better suited for this task and integrate them in our system, QUDE. Via extensive experiments on both real-world and synthetic data sets we demonstrate how QUDE can help experts and novice users alike to efficiently control false discoveries.
Zheguang Zhao, Lorenzo De Stefani, Emanuel Zgraggen, Carsten Binnig, Eli Upfal, Tim Kraska
SIGMOD Conference3
2017 Safe Visual Data Exploration
abstract
Exploring data via visualization has become a popular way to understand complex data. Features or patterns in visualization can be perceived as relevant insights by users, even though they may actually arise from random noise. Moreover, interactive data exploration and visualization recommendation tools can examine a large number of observations, and therefore result in further increasing chance of spurious insights. Thus without proper statistical control, the risk of false discovery renders visual data exploration unsafe and makes users susceptible to questionable inference.To address these problems, we present QUDE, a visual data exploration system that interacts with users to formulate hypotheses based on visualizations and provides interactive control of false discoveries.
Zheguang Zhao, Emanuel Zgraggen, Lorenzo De Stefani, Carsten Binnig, Eli Upfal, Tim Kraska
SIGMOD Conference2
2017 Revisiting Reuse for Approximate Query Processing
abstract
Visual data exploration tools allow users to quickly gather insights from new datasets. As dataset sizes continue to increase, though, new techniques will be necessary to maintain the interactivity guarantees that these tools require. Approximate query processing (AQP) attempts to tackle this problem and allows systems to return query results at "human speed." However, existing AQP techniques start to break down when confronted with ad hoc queries that target the tails of the distribution. We therefore present an AQP formulation that can provide low-error approximate results at interactive speeds, even for queries over rare subpopulations. In particular, our formulation treats query results as random variables in order to leverage the ample opportunities for result reuse inherent in interactive data exploration. As part of our approach, we apply a variety of optimization techniques that are based on probability theory, including new query rewrite rules and index structures. We implemented these techniques in a prototype system and show that they can achieve interactivity where alternative approaches cannot.
Alex Galakatos, Andrew Crotty, Emanuel Zgraggen, Carsten Binnig, Tim Kraska
Proc. VLDB Endow.3
2017 How Progressive Visualizations Affect Exploratory Analysis
abstract
The stated goal for visual data exploration is to operate at a rate that matches the pace of human data analysts, but the ever increasing amount of data has led to a fundamental problem: datasets are often too large to process within interactive time frames. Progressive analytics and visualizations have been proposed as potential solutions to this issue. By processing data incrementally in small chunks, progressive systems provide approximate query answers at interactive speeds that are then refined over time with increasing precision. We study how progressive visualizations affect users in exploratory settings in an experiment where we capture user behavior and knowledge discovery through interaction logs and think-aloud protocols. Our experiment includes three visualization conditions and different simulated dataset sizes. The visualization conditions are: (1) blocking, where results are displayed only after the entire dataset has been processed; (2) instantaneous, a hypothetical condition where results are shown almost immediately; and (3) progressive, where approximate results are displayed quickly and then refined over time. We analyze the data collected in our experiment and observe that users perform equally well with either instantaneous or progressive visualizations in key metrics, such as insight discovery rates and dataset coverage, while blocking visualizations have detrimental effects.
Emanuel Zgraggen, Alex Galakatos, Andrew Crotty, Jean-Daniel Fekete, Tim Kraska
IEEE Trans. Vis. Comput. Graph.1
2015 (s|qu)eries: Visual Regular Expressions for Querying and Exploring Event Sequences
abstract
Many different domains collect event sequence data and rely on finding and analyzing patterns within it to gain meaningful insights. Current systems that support such queries either provide limited expressiveness, hinder exploratory workflows or present interaction and visualization models which do not scale well to large and multi-faceted data sets. In this paper we present (s|qu)eries (pronounced "Squeries"), a visual query interface for creating queries on sequences (series) of data, based on regular expressions. (s|qu)eries is a touch-based system that exposes the full expressive power of regular expressions in an approachable way and interleaves query specification with result visualizations. Being able to visually investigate the results of different query-parts supports debugging and encourages iterative query-building as well as exploratory work-flows. We validate our design and implementation through a set of informal interviews with data scientists that analyze event sequences on a daily basis.
Emanuel Zgraggen, Steven Mark Drucker, Danyel Fisher, Robert DeLine
CHI1
2015 Evaluating Subjective Accuracy in Time Series Pattern-Matching Using Human-Annotated Rankings
abstract
Finding patterns is a common task in time series analysis which has gained a lot of attention across many fields. A multitude of similarity measures have been introduced to perform pattern searches. The accuracy of such measures is often evaluated objectively using a one nearest neighbor classification (1NN) on labeled time series or through clustering. Prior work often disregards the subjective similarity of time series which can be pivotal in systems where a user specified pattern is used as input and a similarity-based ranking is expected as output (query-by-example). In this paper, we describe how a human-annotated ranking based on real-world queries and datasets can be created using simple crowdsourcing tasks and use this ranking as ground-truth to evaluate the perceived accuracy of existing time series similarity measures. Furthermore, we show how different sampling strategies and time series representations of pen-drawn queries effect the precision of these similarity measures and provide a publicly available dataset which can be used to optimize existing and future similarity search algorithms.
Philipp Eichmann, Emanuel Zgraggen
IUI2
2015 Vizdom: Interactive Analytics through Pen and Touch
abstract
Machine learning (ML) and advanced statistics are important tools for drawing insights from large datasets. However, these techniques often require human intervention to steer computation towards meaningful results. In this demo, we present V izdom , a new system for interactive analytics through pen and touch. V izdom 's frontend allows users to visually compose complex workflows of ML and statistics operators on an interactive whiteboard, and the back-end leverages recent advances in workflow compilation techniques to run these computations at interactive speeds. Additionally, we are exploring approximation techniques for quickly visualizing partial results that incrementally refine over time. This demo will show V izdom 's capabilities by allowing users to interactively build complex analytics workflows using real-world datasets.
Andrew Crotty, Alex Galakatos, Emanuel Zgraggen, Carsten Binnig, Tim Kraska
Proc. VLDB Endow.3
2014 PanoramicData: Data Analysis through Pen & Touch
abstract
Interactively exploring multidimensional datasets requires frequent switching among a range of distinct but inter-related tasks (e.g., producing different visuals based on different column sets, calculating new variables, and observing the interactions between sets of data). Existing approaches either target specific different problem domains (e.g., data-transformation or data-presentation) or expose only limited aspects of the general exploratory process; in either case, users are forced to adopt coping strategies (e.g., arranging windows or using undo as a mechanism for comparison instead of using side-by-side displays) to compensate for the lack of an integrated suite of exploratory tools. PanoramicData (PD) addresses these problems by unifying a comprehensive set of tools for visual data exploration into a hybrid pen and touch system designed to exploit the visualization advantages of large interactive displays. PD goes beyond just familiar visualizations by including direct UI support for data transformation and aggregation, filtering and brushing. Leveraging an unbounded whiteboard metaphor, users can combine these tools like building blocks to create detailed interactive visual display networks in which each visualization can act as a filter for others. Further, by operating directly on relational-databases, PD provides an approachable visual language that exposes a broad set of the expressive power of SQL including functionally complete logic filtering, computation of aggregates and natural table joins. To understand the implications of this novel approach, we conducted a formative user study with both data and visualization experts. The results indicated that the system provided a fluid and natural user experience for probing multi-dimensional data and was able to cover the full range of queries that the users wanted to pose.
Emanuel Zgraggen, Robert C. Zeleznik, Steven Mark Drucker
IEEE Trans. Vis. Comput. Graph.1