Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sven Langenecker

dblp:287/7567 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2024
0009-0002-2809-5331ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 56% Data integration and cleaning · 28% Distributed and cloud data management · 8%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management › weak supervision
data programming
0.712023
Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023
Data integration and cleaning › table understanding › table annotation
semantic type detection
0.712023
Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023
Machine learning and data management › training data management
training data generation
0.712023
Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023
Distributed and cloud data management
data lake
0.212023
Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023
Knowledge graphs › semantic web
semantic annotation
0.212023
Steered Training Data Generation for Learned Semantic Type Detection · Proc. ACM Manag. Data 2023

Methods — techniques the papers use, named apart from their topics

steered-labeling · 0.7fine-tuning · 0.7data programming · 0.7
YearPublicationVenuePosition
2024 Pythagoras: Semantic Type Detection of Numerical Data in Enterprise Data Lakes
Sven Langenecker, Christoph Sturm, Christian Schalles, Carsten Binnig
EDBT1
2023 Steered Training Data Generation for Learned Semantic Type Detection
abstract
In this paper, we introduce STEER to adapt learned semantic type extraction approaches to a new, unseen data lake. STEER provides a data programming framework for semantic labeling which is used to generate new labeled training data with minimal overhead. At its core, STEER comes with a novel training data generation procedure called Steered-Labeling that can generate high quality training data not only for non-numeric but also for numerical columns. With this generated training data STEER is able to fine-tune existing learned semantic type extraction models. We evaluate our approach on four different data lakes and show that we can significantly improve the performance of two different types of learned models across all data lakes.
Sven Langenecker, Christoph Sturm, Christian Schalles, Carsten Binnig
Proc. ACM Manag. Data1