Sebastian Eggers

dblp:415/9353 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0005-6076-7101ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Program analysis · 67% Requirements engineering and software design · 33%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
data flow analysis
0.912025
APEX-DAG: Library and Language independent Pipeline EXtraction · Proc. VLDB Endow. 2025
Requirements engineering and software design › software architecture › software architecture analysis
software architecture recovery
0.912025
APEX-DAG: Library and Language independent Pipeline EXtraction · Proc. VLDB Endow. 2025
Program analysis
static analysis
0.912025
APEX-DAG: Library and Language independent Pipeline EXtraction · Proc. VLDB Endow. 2025
Machine learning and data management
machine learning pipeline
0.312025
APEX-DAG: Library and Language independent Pipeline EXtraction · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

static code analysis · 1.7graph attention network · 1.7
YearPublicationVenuePosition
2025 APEX-DAG: Library and Language independent Pipeline EXtraction
abstract
Modern data-driven systems often rely on complex pipelines to process and transform data for downstream machine learning (ML) tasks. Extracting these pipelines and understanding their structure is critical for ensuring transparency, performance optimization, and maintainability, especially in large-scale projects. In this work, we introduce a novel system, APEX-DAG ( A utomating P ipeline EX traction with D ataflow, Static Code A nalysis, and G raph Attention Networks), which automates the extraction of data pipelines from computational notebooks or scripts. Unlike execution-based methods, APEX-DAG leverages static code analysis to identify the dataflow, transformations, and dependencies within ML workflows without executing the code or the need to alter the code. Further, after an initial training phase, our system can identify pipelines that built with previously unseen libraries.
Sebastian Eggers, Nina Zukowska, Ziawasch Abedjan
Proc. VLDB Endow.1