EDBT 2026 Demo / reviewers in the wild / expert
Sven Langenecker
dblp:287/7567
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2024
0009-0002-2809-5331ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 56% Data integration and cleaning · 28% Distributed and cloud data management · 8% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management › weak supervision
data programming |
0.7 | 1 | 2023 | Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023 |
Data integration and cleaning › table understanding › table annotation
semantic type detection |
0.7 | 1 | 2023 | Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023 |
Machine learning and data management › training data management
training data generation |
0.7 | 1 | 2023 | Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023 |
Distributed and cloud data management
data lake |
0.2 | 1 | 2023 | Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023 |
Knowledge graphs › semantic web
semantic annotation |
0.2 | 1 | 2023 | Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023 |
Methods — techniques the papers use, named apart from their topics
steered-labeling · 0.7fine-tuning · 0.7data programming · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Pythagoras: Semantic Type Detection of Numerical Data in Enterprise Data Lakes
Sven Langenecker, Christoph Sturm, Christian Schalles, Carsten Binnig |
EDBT | 1 |
| 2023 | Steered Training Data Generation for Learned Semantic Type DetectionabstractIn this paper, we introduce STEER to adapt learned semantic type extraction approaches to a new, unseen data lake. STEER provides a data programming framework for semantic labeling which is used to generate new labeled training data with minimal overhead. At its core, STEER comes with a novel training data generation procedure called Steered-Labeling that can generate high quality training data not only for non-numeric but also for numerical columns. With this generated training data STEER is able to fine-tune existing learned semantic type extraction models. We evaluate our approach on four different data lakes and show that we can significantly improve the performance of two different types of learned models across all data lakes. Sven Langenecker, Christoph Sturm, Christian Schalles, Carsten Binnig |
Proc. ACM Manag. Data | 1 |