EDBT 2026 Demo / reviewers in the wild / expert
Dan Sotolongo
dblp:300/3960
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2023
0009-0006-0646-0958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 57% Data stream processing · 43% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
incremental computation |
0.7 | 1 | 2023 | What's the Difference? Incremental Processing with Change Queries in Snowflake · Proc. ACM Manag. Data 2023 |
Query processing and optimization › view maintenance
incremental view maintenance |
0.7 | 1 | 2023 | What's the Difference? Incremental Processing with Change Queries in Snowflake · Proc. ACM Manag. Data 2023 |
Data stream processing
stream processing systems |
0.5 | 1 | 2021 | Watermarks in Stream Processing Systems: Semantics and Comparative Analysis of Apache Flink and Google Cloud Dataflow · Proc. VLDB Endow. 2021 |
Methods — techniques the papers use, named apart from their topics
stream objects · 0.7DML · 0.7cubic spline models · 0.5comparative analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | What's the Difference? Incremental Processing with Change Queries in SnowflakeabstractIncremental algorithms are the heart and soul of stream processing. Low latency results depend on the ability to react to the subset of changes in a dataset over time rather than reprocessing the entirety of a dataset as it evolves. But while the SQL language is well suited for representing streams of changes (via tables) and their application to tables over time (via DML), it entirely lacks a method to query the changes to a table or view in the first place. In this paper, we present CHANGES queries and STREAM objects, Snowflake's primitives for querying and consuming incremental changes to table objects over time. CHANGES queries and STREAMs have been in use within Snowflake for three years, and see broad adoption across our customers. We describe the semantics of these primitives, discuss the implementation challenges, present an analysis of their usage at Snowflake, and contrast with other offerings. Tyler Akidau, Paul Barbier, Istvan Cseri, Fabian Hueske, Tyler Jones, Sasha Lionheart, Daniel Mills, Dzmitry Pauliukevich, Lukas Probst, Niklas Semmler, Dan Sotolongo, Boyuan Zhang 0004 |
Proc. ACM Manag. Data | 11 |
| 2021 | Watermarks in Stream Processing Systems: Semantics and Comparative Analysis of Apache Flink and Google Cloud DataflowabstractStreaming data processing is an exercise in taming disorder: from oftentimes huge torrents of information, we hope to extract powerful and timely analyses. But when dealing with streaming data, the unbounded and temporally disordered nature of real-world streams introduces a critical challenge: how does one reason about the completeness of a stream that never ends? In this paper, we present a comprehensive definition and analysis of watermarks , a key tool for reasoning about temporal completeness in infinite streams. First, we describe what watermarks are and why they are important, highlighting how they address a suite of stream processing needs that are poorly served by eventually-consistent approaches: • Computing a single correct answer, as in notifications. • Reasoning about a lack of data, as in dip detection. • Performing non-incremental processing over temporal subsets of an infinite stream, as in statistical anomaly detection with cubic spline models. • Safely and punctually garbage collecting obsolete inputs and intermediate state. • Surfacing a reliable signal of overall pipeline health . Second, we describe, evaluate, and compare the semantically equivalent, but starkly different, watermark implementations in two modern stream processing engines: Apache Flink and Google Cloud Dataflow. Edmon Begoli, Tyler Akidau, Slava Chernyak, Fabian Hueske, Kathryn Knight, Kenneth L. Knowles, Daniel Mills, Dan Sotolongo |
Proc. VLDB Endow. | 8 |