EDBT 2026 Demo / reviewers in the wild / expert
Georgi Krastev
dblp:96/1799
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0002-2904-3317ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 67% Hardware accelerators and domain-specific architectures · 33% | |
| Software engineering, system software, and programming languages
2 papers |
Programming languages and type systems · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Distributed and cloud data management · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Programming languages and type systems
domain-specific languages |
0.4 | 1 | 2019 | Representations and Optimizations for Embedded Parallel Dataflow Languages · ACM Trans. Database Syst. 2019 |
Hardware accelerators and domain-specific architectures
dataflow optimization |
0.4 | 1 | 2019 | Representations and Optimizations for Embedded Parallel Dataflow Languages · ACM Trans. Database Syst. 2019 |
Parallel and multicore computing › programming models
embedded domain-specific language |
0.4 | 1 | 2019 | Representations and Optimizations for Embedded Parallel Dataflow Languages · ACM Trans. Database Syst. 2019 |
Parallel and multicore computing › parallel computation models
parallel dataflow |
0.4 | 1 | 2019 | Representations and Optimizations for Embedded Parallel Dataflow Languages · ACM Trans. Database Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
structural recursion · 0.8bag algebra · 0.8quasi-quotation · 0.5deep language embedding · 0.5monads · 0.4monad · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Representations and Optimizations for Embedded Parallel Dataflow LanguagesabstractParallel dataflow engines such as Apache Hadoop, Apache Spark, and Apache Flink are an established alternative to relational databases for modern data analysis applications. A characteristic of these systems is a scalable programming model based on distributed collections and parallel transformations expressed by means of second-order functions such as map and reduce. Notable examples are Flink’s DataSet and Spark’s RDD programming abstractions. These programming models are realized as EDSLs—domain specific languages embedded in a general-purpose host language such as Java, Scala, or Python. This approach has several advantages over traditional external DSLs such as SQL or XQuery. First, syntactic constructs from the host language (e.g., anonymous functions syntax, value definitions, and fluent syntax via method chaining) can be reused in the EDSL. This eases the learning curve for developers already familiar with the host language. Second, it allows for seamless integration of library methods written in the host language via the function parameters passed to the parallel dataflow operators. This reduces the effort for developing analytics dataflows that go beyond pure SQL and require domain-specific logic. At the same time, however, state-of-the-art parallel dataflow EDSLs exhibit a number of shortcomings. First, one of the main advantages of an external DSL such as SQL—the high-level, declarative Select-From-Where syntax—is either lost completely or mimicked in a non-standard way. Second, execution aspects such as caching, join order, and partial aggregation have to be decided by the programmer. Optimizing them automatically is very difficult due to the limited program context available in the intermediate representation of the DSL. In this article, we argue that the limitations listed above are a side effect of the adopted type-based embedding approach. As a solution, we propose an alternative EDSL design based on quotations. We present a DSL embedded in Scala and discuss its compiler pipeline, intermediate representation, and some of the enabled optimizations. We promote the algebraic type of bags in union representation as a model for distributed collections and its associated structural recursion scheme and monad as a model for parallel collection processing. At the source code level, Scala’s comprehension syntax over a bag monad can be used to encode Select-From-Where expressions in a standard way. At the intermediate representation level, maintaining comprehensions as a first-class citizen can be used to simplify the design and implementation of holistic dataflow optimizations that accommodate for nesting and control-flow. The proposed DSL design therefore reconciles the benefits of embedded parallel dataflow DSLs with the declarativity and optimization potential of external DSLs like SQL. Alexander Alexandrov 0001, Georgi Krastev, Volker Markl |
ACM Trans. Database Syst. | 2 |
| 2016 | Emma in Action: Declarative Dataflows for Scalable Data AnalysisabstractParallel dataflow APIs based on second-order functions were originally seen as a flexible alternative to SQL. Over time, however, their complexity increased due to the number of physical aspects that had to be exposed by the underlying engines in order to facilitate efficient execution. To retain a sufficient level of abstraction and lower the barrier of entry for data scientists, projects like Spark and Flink currently offer domain-specific APIs on top of their parallel collection abstractions. This demonstration highlights the benefits of an alternative design based on deep language embedding. We showcase Emma - a programming language embedded in Scala. Emma promotes parallel collection processing through native constructs like Scala's for-comprehensions - a declarative syntax akin to SQL. In addition, Emma also advocates quasi-quoting the entire data analysis algorithm rather than its individual dataflow expressions. This allows for decomposing the quoted code into (sequential) control flow and (parallel) dataflow fragments, optimizing the dataflows in context, and transparently offloading them to an engine like Spark or Flink. The proposed design promises increased programmer productivity due to avoiding an impedance mismatch, thereby reducing the lag times and cost of data analysis. Alexander Alexandrov 0001, Andreas Salzmann, Georgi Krastev, Asterios Katsifodimos, Volker Markl |
SIGMOD Conference | 3 |