Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dustin Lange

dblp:25/8701 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 5 first-authorArtificial intelligence and machine learning · 5 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Data integration and cleaning · 65% Data stream processing · 13% Machine learning and data management · 10%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 50% Distributed systems · 50%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data quality
0.522019
DataWig: Missing Value Imputation for Tables · J. Mach. Learn. Res. 2019
Automating Large-Scale Data Quality Verification · Proc. VLDB Endow. 2018
Data stream processing
incremental validation
0.412019
Unit Testing Data with Deequ · SIGMOD Conference 2019
Data integration and cleaning › missing data
missing value imputation
0.412019
DataWig: Missing Value Imputation for Tables · J. Mach. Learn. Res. 2019
Data mining › predictive modeling › forecasting
demand prediction
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Machine learning and data management
machine learning pipeline
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
High-performance computing
cluster computing
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Distributed systems
distributed machine learning
0.312017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017
Query processing and optimization
aggregate query processing
0.112018
Automating Large-Scale Data Quality Verification · Proc. VLDB Endow. 2018
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting
0.112017
Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017

Methods — techniques the papers use, named apart from their topics

ensembling · 0.9feature engineering · 0.6monoid · 0.4hyperparameter tuning · 0.4deep learning feature extraction · 0.4declarative constraint API · 0.4apache spark · 0.4anomaly detection on quality metrics · 0.4algebraic state · 0.4machine learning · 0.3constraint suggestion · 0.3
YearPublicationVenuePosition
2019 Differential Data Quality Verification on Partitioned Data
abstract
Modern companies and institutions rely on data to guide every single decision. Missing or incorrect information seriously compromises any decision process. In previous work, we presented Deequ, a Spark-based library for automating the verification of data quality at scale. Deequ provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables "unit tests for data". However, we found that the previous computational model of Deequ is not flexible enough for many scenarios in modern data pipelines, which handle large, partitioned datasets. Such scenarios require the evaluation of dataset-level quality constraints after individual partition updates, without having to re-read already processed partitions. Additionally, such scenarios often require the verification of data quality on select combinations of partitions. We therefore present a differential generalization of the computational model of Deequ, based on algebraic states with monoid properties. We detail how to efficiently implement the corresponding operators and aggregation functions in Apache Spark. Furthermore, we show how to optimize the resulting workloads to minimize the required number of passes over the data, and empirically validate that our approach decreases the runtimes for updating data metrics under data changes and for different combinations of partitions.
Sebastian Schelter, Stefan Grafberger, Philipp Schmidt 0002, Tammo Rukat, Mario Kießling, Andrey Taptunov, Felix Bießmann, Dustin Lange
ICDE8
2019 Unit Testing Data with Deequ
abstract
Modern companies and institutions rely on data to guide every single decision. Missing or incorrect information seriously compromises any decision process. We demonstrate "Deequ", an Apache Spark-based library for automating the verification of data quality at scale. This library provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables "unit tests for data". Deequ is available as open source, meets the requirements of production use cases at Amazon, and scales to datasets with billions of records if the constraints to evaluate are chosen carefully. Our demonstration walks attendees through a fictitious business use case of validating daily product reviews from a public dataset, and is executed in a proprietary interactive notebook environment. We show attendees how to define data unit tests from automatically suggested constraints and how to create customized tests. Additionally, we demonstrate how to apply Deequ to validate incrementally growing datasets, and give examples of how to configure anomaly detection algorithms on time series of data quality metrics to further automate the data validation.
Sebastian Schelter, Felix Bießmann, Dustin Lange, Tammo Rukat, Philipp Schmidt 0002, Stephan Seufert, Pierre Brunelle, Andrey Taptunov
SIGMOD Conference3
2019 DataWig: Missing Value Imputation for Tables
abstract
With the growing importance of machine learning (ML) algorithms for practical applications, reducing data quality problems in ML pipelines has become a major focus of research. In many cases missing values can break data pipelines which makes completeness one of the most impactful data quality challenges. Current missing value imputation methods are focusing on numerical or categorical data and can be difficult to scale to datasets with millions of rows. We release DataWig, a robust and scalable approach for missing value imputation that can be applied to tables with heterogeneous data types, including unstructured text. DataWig combines deep learning feature extractors with automatic hyperparameter tuning. This enables users without a machine learning background, such as data engineers, to impute missing values with minimal effort in tables with more heterogeneous data types than supported in existing libraries, while requiring less glue code for feature engineering and offering more flexible modelling options. We demonstrate that DataWig compares favourably to existing imputation packages. Source code, documentation, and unit tests for this package are available at: https://github.com/awslabs/datawig
Felix Bießmann, Tammo Rukat, Philipp Schmidt 0002, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, David Salinas
J. Mach. Learn. Res.7
2018 "Deep" Learning for Missing Value Imputationin Tables with Non-Numerical Data
abstract
The success of applications that process data critically depends on the quality of the ingested data. Completeness of a data source is essential in many cases. Yet, most missing value imputation approaches suffer from severe limitations. They are almost exclusively restricted to numerical data, and they either offer only simple imputation methods or are difficult to scale and maintain in production. Here we present a robust and scalable approach to imputation that extends to tables with non-numerical values, including unstructured text data in diverse languages. Experiments on public data sets as well as data sets sampled from a large product catalog in different languages (English and Japanese) demonstrate that the proposed approach is both scalable and yields more accurate imputations than previous approaches. Training on data sets with several million rows is a matter of minutes on a single machine. With a median imputation F1 score of 0.93 across a broad selection of data sets our approach achieves on average a 23-fold improvement compared to mode imputation. While our system allows users to apply state-of-the-art deep learning models if needed, we find that often simple linear n-gram models perform on par with deep learning methods at a much lower operational cost. The proposed method learns all parameters of the entire imputation pipeline automatically in an end-to-end fashion, rendering it attractive as a generic plugin both for engineers in charge of data pipelines where data completeness is relevant, as well as for practitioners without expertise in machine learning who need to impute missing values in tables with non-numerical data.
Felix Bießmann, David Salinas, Sebastian Schelter, Philipp Schmidt 0002, Dustin Lange
CIKM5
2018 Automating Large-Scale Data Quality Verification
abstract
Modern companies and institutions rely on data to guide every single business process and decision. Missing or incorrect information seriously compromises any decision process downstream. Therefore, a crucial, but tedious task for everyone involved in data processing is to verify the quality of their data. We present a system for automating the verification of data quality at scale, which meets the requirements of production use cases. Our system provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables 'unit tests' for data. We efficiently execute the resulting constraint validation workload by translating it to aggregation queries on Apache Spark. Our platform supports the incremental validation of data quality on growing datasets, and leverages machine learning, e.g., for enhancing constraint suggestions, for estimating the 'predictability' of a column, and for detecting anomalies in historic data quality time series. We discuss our design decisions, describe the resulting system architecture, and present an experimental evaluation on various datasets.
Sebastian Schelter, Dustin Lange, Philipp Schmidt 0002, Meltem Celikel, Felix Bießmann, Andreas Grafberger
Proc. VLDB Endow.2
2017 Probabilistic Demand Forecasting at Scale
abstract
We present a platform built on large-scale, data-centric machine learning (ML) approaches, whose particular focus is demand forecasting in retail. At its core, this platform enables the training and application of probabilistic demand forecasting models, and provides convenient abstractions and support functionality for forecasting problems. The platform comprises of a complex end-to-end machine learning system built on Apache Spark, which includes data preprocessing, feature engineering, distributed learning, as well as evaluation, experimentation and ensembling. Furthermore, it meets the demands of a production system and scales to large catalogues containing millions of items. We describe the challenges of building such a platform and discuss our design decisions. We detail aspects on several levels of the system, such as a set of general distributed learning schemes, our machinery for ensembling predictions, and a high-level dataflow abstraction for modeling complex ML pipelines. To the best of our knowledge, we are not aware of prior work on real-world demand forecasting systems which rivals our approach in terms of scalability.
Joos-Hendrik Böse, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Dustin Lange, David Salinas, Sebastian Schelter, Matthias W. Seeger, Yuyang Wang 0001
Proc. VLDB Endow.5
2013 Bulk sorted access for efficient top-k retrieval
abstract
Efficient top-k retrieval of records from a database has been an active research field for many years. We approach the problem from a real-world application point of view, in which the order of records according to some similarity function on an attribute is not unique: Many records have same values in several attributes and thus their ranking in those attributes is arbitrary. For instance, in large person databases many individuals have the same first name, the same date of birth, or live in the same city. Existing algorithms, such as the Threshold Algorithm (TA), are ill-equipped to handle such cases efficiently.
Dustin Lange, Felix Naumann
SSDBM1
2013 Cost-aware query planning for similarity search
Dustin Lange, Felix Naumann
Inf. Syst.1
2013 Cross-lingual entity matching and infobox alignment in Wikipedia
Daniel Rinser, Dustin Lange, Felix Naumann
Inf. Syst.2
2012 Efficient Similarity Search in Very Large String Sets
Dandy Fenz, Dustin Lange, Astrid Rheinländer, Felix Naumann, Ulf Leser
SSDBM2
2011 Frequency-aware similarity measures: why Arnold Schwarzenegger is always a duplicate
abstract
Measuring the similarity of two records is a challenging problem, but necessary for fundamental tasks, such as duplicate detection and similarity search. By exploiting frequencies of attribute values, many similarity measures can be improved: In a person table with U.S. citizens, Arnold Schwarzenegger is a very rare name. If we find several Arnold Schwarzeneggers in it, it is very likely that these are duplicates. We are then less strict when comparing other attribute values, such as birth date or address. We put this intuition to use by partitioning compared record pairs according to frequencies of attribute values. For example, we could create three partitions from our data: Partition 1 contains all pairs with rare names, Partition 2 all pairs with medium frequent names, and Partition 3 all pairs with frequent names. For each partition, we learn a different similarity measure: we apply machine learning techniques to combine a set of base similarity measures into an overall measure. To determine a good partitioning, we compare different partitioning strategies. We achieved best results with a novel algorithm inspired by genetic programming.
Dustin Lange, Felix Naumann
CIKM1
2011 Efficient similarity search: arbitrary similarity measures, arbitrary composition
abstract
Given a (large) set of objects and a query, similarity search aims to find all objects similar to the query. A frequent approach is to define a set of base similarity measures for the different aspects of the objects, and to build light-weight similarity indexes on these measures. To determine the overall similarity of two objects, the results of these base measures are composed, e.g., using simple aggregates or more involved machine learning techniques. We propose the first solution to this search problem that does not place any restrictions on the similarity measures, the composition technique, or the data set size. We define the query plan optimization problem to determine the best query plan using the similarity indexes. A query plan must choose which individual indexes to access and which thresholds to apply. The plan result should be as complete as possible within some cost threshold. We propose the approximative top neighborhood algorithm, which determines a near-optimal plan while significantly reducing the amount of candidate plans to be considered. An exact version of the algorithm determines the optimal solution. Evaluation on real-world data indicates that both versions clearly outperform a complete search of the query plan space.
Dustin Lange, Felix Naumann
CIKM1
2010 Extracting structured information from Wikipedia articles to populate infoboxes
abstract
Roughly every third Wikipedia article contains an infobox - a table that displays important facts about the subject in attribute-value form. The schema of an infobox, i.e., the attributes that can be expressed for a concept, is defined by an infobox template. Often, authors do not specify all template attributes, resulting in incomplete infoboxes. With iPopulator, we introduce a system that automatically populates infoboxes of Wikipedia articles by extracting attribute values from the article's text. In contrast to prior work, iPopulator detects and exploits the structure of attribute values for independently extracting value parts. We have tested iPopulator on the entire set of infobox templates and provide a detailed analysis of its effectiveness. For instance, we achieve an average extraction precision of 91% for 1,727 distinct infobox template attributes.
Dustin Lange, Christoph Böhm 0001, Felix Naumann
CIKM1