Philipp Eichmann

dblp:160/4301 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 54% Database system architecture and tuning · 19% Machine learning and data management · 12%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.512021
Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021
Query processing and optimization › approximate query processing
sampling-based approximate query processing
0.512021
Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021
Database system architecture and tuning
database benchmarking
0.412020
IDEBench: A Benchmark for Interactive Data Exploration · SIGMOD Conference 2020
Query processing and optimization
interactive data exploration
0.412020
IDEBench: A Benchmark for Interactive Data Exploration · SIGMOD Conference 2020
Performance modeling and evaluation
benchmarking
0.412020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Performance modeling and evaluation › benchmarking
database system benchmarking
0.412020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Data integration and cleaning
data preprocessing
0.412019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019
Machine learning and data management
machine learning pipeline
0.412019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019
Visualization and visual analytics › visual analytics
anomaly detection visualization
0.412019
Visual Exploration of Time Series Anomalies with Metro-Viz · SIGMOD Conference 2019
Visualization and visual analytics
time series visualization
0.412019
Visual Exploration of Time Series Anomalies with Metro-Viz · SIGMOD Conference 2019
Query processing and optimization
interactive query workload
0.112020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Visualization and visual analytics
interactive data exploration
0.112020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Data mining
anomaly detection
0.112019
Visual Exploration of Time Series Anomalies with Metro-Viz · SIGMOD Conference 2019
Human-AI interaction
interactive machine learning
0.112019
Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019

Methods — techniques the papers use, named apart from their topics

model selection · 0.8hyperparameter tuning · 0.8sampling · 0.5progressive computation · 0.5user study · 0.4data-management strategies · 0.4data management strategies · 0.4
YearPublicationVenuePosition
2021 Davos: A System for Interactive Data-Driven Decision Making
abstract
Recently, a new horizon in data analytics, prescriptive analytics, is becoming more and more important to make data-driven decisions. As opposed to the progress of democratizing data acquisition and access, making data-driven decisions remains a significant challenge for people without technical expertise. In this regard, existing tools for data analytics which were designed decades ago still present a high bar for domain experts, and removing this bar requires a fundamental rethinking of both interface and backend. At Einblick, an MIT/Brown spin-off based on the Northstar project, we have been building the next generation analytics tool in the last few years. To overcome the shortcomings of existing processing engines, we propose Davos , Einblick's novel backend. Davos combines aspects of progressive computation, approximate query processing and sampling, with a specific focus on supporting user-defined operations. Moreover, Davos optimizes multi-tenant scenarios to promote collaboration. Both empirical evaluation and user study verify that Davos can greatly empower data analytics for new needs.
Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Philipp Eichmann, Navid Karimeddiny, Charlie Meyer, Wesley Runnels, Tim Kraska
Proc. VLDB Endow.4
2020 Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data
abstract
In this paper, we present a new benchmark to validate the suitability of database systems for interactive visualization workloads. While there exist proposals for evaluating database systems on interactive data exploration workloads, none rely on real user traces for database benchmarking. To this end, our long term goal is to collect user traces that represent workloads with different exploration characteristics. In this paper, we present an initial benchmark that focuses on "crossfilter"-style applications, which are a popular interaction type for data exploration and a particularly demanding scenario for testing database system performance. We make our benchmark materials, including input datasets, interaction sequences, corresponding SQL queries, and analysis code, freely available as a community resource, to foster further research in this area: https://osf.io/9xerb/?view_only=81de1a3f99d04529b6b173a3bd5b4d23.
Leilani Battle, Philipp Eichmann, Marco Angelini, Tiziana Catarci, Giuseppe Santucci, Yukun Zheng, Carsten Binnig, Jean-Daniel Fekete, Dominik Moritz
SIGMOD Conference2
2020 IDEBench: A Benchmark for Interactive Data Exploration
abstract
In recent years, many query processing techniques have been developed to better support interactive data exploration (IDE) of large structured datasets. To evaluate and compare database engines in terms of how well they support such workloads, experimenters have mostly used self-designed evaluation procedures rather than established benchmarks. In this paper we argue that this is due to the fact that the workloads and metrics of popular analytical benchmarks such as TPC-H or TPC-DS were designed for traditional performance reporting scenarios, and do not capture distinctive IDE characteristics. Guided by the findings of several user studies we present a new benchmark called IDEBench, designed to evaluate database engines based on common IDE workflows and metrics that matter to the end-user. We demonstrate the applicability of IDEBench through a number of experiments with five different database engines, and present and discuss our findings.
Philipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim Kraska
SIGMOD Conference1
2020 Orchard: Exploring Multivariate Heterogeneous Networks on Mobile Phones
abstract
Abstract People are becoming increasingly sophisticated in their ability to navigate information spaces using search, hyperlinks, and visualization. But, mobile phones preclude the use of multiple coordinated views that have proven effective in the desktop environment (e.g., for business intelligence or visual analytics). In this work, we propose to model information as multivariate heterogeneous networks to enable greater analytic expression for a range of sensemaking tasks while suggesting a new, list‐based paradigm with gestural navigation of structured information spaces on mobile phones. We also present a mobile application, called Orchard, which combines ideas from both faceted search and interactive network exploration in a visual query language to allow users to collect facets of interest during exploratory navigation. Our study showed that users could collect and combine these facets with Orchard, specifying network queries and projections that would only have been possible previously using complex data tools or custom data science.
Philipp Eichmann, Darren Edge, Nathan Evans, Bongshin Lee, Matthew Brehmer, Christopher M. White
Comput. Graph. Forum1
2019 Visual Exploration of Time Series Anomalies with Metro-Viz
abstract
This demo presents a novel data visualization solution for exploring the results of time series anomaly detection systems. When anomalies are reported, there is a need to reason about the results. We introduce Metro-Viz -- a visual tool to assist data scientists in performing this analysis. Metro-Viz offers a rich set of interaction features (e.g., comparative analysis, what-if testing) backed by data management strategies specifically tailored to the workload. We show our tool in action via multiple time series datasets and anomaly detectors.
Philipp Eichmann, Franco Solleza, Nesime Tatbul, Stanley B. Zdonik
SIGMOD Conference1
2019 Democratizing Data Science through Interactive Curation of ML Pipelines
abstract
Statistical knowledge and domain expertise are key to extract actionable insights out of data, yet such skills rarely coexist together. In Machine Learning, high-quality results are only attainable via mindful data preprocessing, hyperparameter tuning and model selection. Domain experts are often overwhelmed by such complexity, de-facto inhibiting a wider adoption of ML techniques in other fields. Existing libraries that claim to solve this problem, still require well-trained practitioners. Those frameworks involve heavy data preparation steps and are often too slow for interactive feedback from the user, severely limiting the scope of such systems.
Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Ferdinand Kossmann, Philipp Eichmann, Yeounoh Chung, Carsten Binnig, Eli Upfal, Tim Kraska
SIGMOD Conference5
2017 NuSys: Towards a Document IDE for Knowledge Work
abstract
Knowledge workers consume and annotate digital documents such as PDF files, videos, images and text notes - in some cases collaboratively - to form mental models and gain insight. An abundance of software solutions and utilities that were designed to assist users in stages of this process but not in the process as a whole, which makes knowledge work with documents unnecessarily inefficient. In this paper, we introduce ideas on how to streamline common knowledge worker tasks, such as collaboratively searching, gathering and freely arranging fragments of various media documents to gain understanding and then transforming emergent insights into interactive structured visualizations. Furthermore, we present NuSys, an integrated development environment (IDE) specialized for document-centric workflows, that implements the core of these ideas.
Philipp Eichmann, Trent Green, Robert C. Zeleznik, Andries van Dam
DocEng1
2015 Evaluating Subjective Accuracy in Time Series Pattern-Matching Using Human-Annotated Rankings
abstract
Finding patterns is a common task in time series analysis which has gained a lot of attention across many fields. A multitude of similarity measures have been introduced to perform pattern searches. The accuracy of such measures is often evaluated objectively using a one nearest neighbor classification (1NN) on labeled time series or through clustering. Prior work often disregards the subjective similarity of time series which can be pivotal in systems where a user specified pattern is used as input and a similarity-based ranking is expected as output (query-by-example). In this paper, we describe how a human-annotated ranking based on real-world queries and datasets can be created using simple crowdsourcing tasks and use this ranking as ground-truth to evaluate the perceived accuracy of existing time series similarity measures. Furthermore, we show how different sampling strategies and time series representations of pen-drawn queries effect the precision of these similarity measures and provide a publicly available dataset which can be used to optimize existing and future similarity search algorithms.
Philipp Eichmann, Emanuel Zgraggen
IUI1