EDBT 2026 Demo / reviewers in the wild / expert
Reuven Lax
dblp:135/4663
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data stream processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
fault tolerance |
0.2 | 1 | 2013 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013 |
Data stream processing › stream processing systems
low-latency stream processing |
0.2 | 1 | 2013 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013 |
Distributed systems › fault tolerance
exactly-once processing |
0.2 | 1 | 2013 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013 |
Distributed systems
fault tolerance |
0.2 | 1 | 2013 | MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013 |
Methods — techniques the papers use, named apart from their topics
persistent state management · 0.3directed computation graph · 0.3windowing semantics · 0.2dataflow model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data ProcessingabstractUnbounded, unordered, global-scale datasets are increasingly common in day-to-day business (e.g. Web logs, mobile usage statistics, and sensor networks). At the same time, consumers of these datasets have evolved sophisticated requirements, such as event-time ordering and windowing by features of the data themselves, in addition to an insatiable hunger for faster answers. Meanwhile, practicality dictates that one can never fully optimize along all dimensions of correctness, latency, and cost for these types of input. As a result, data processing practitioners are left with the quandary of how to reconcile the tensions between these seemingly competing propositions, often resulting in disparate implementations and systems. We propose that a fundamental shift of approach is necessary to deal with these evolved requirements in modern data processing. We as a field must stop trying to groom unbounded datasets into finite pools of information that eventually become complete, and instead live and breathe under the assumption that we will never know if or when we have seen all of our data, only that new data will arrive, old data may be retracted, and the only way to make this problem tractable is via principled abstractions that allow the practitioner the choice of appropriate tradeoffs along the axes of interest: correctness, latency, and cost. In this paper, we present one such approach, the Dataflow Model, along with a detailed examination of the semantics it enables, an overview of the core principles that guided its design, and a validation of the model itself via the real-world experiences that led to its development. Tyler Akidau, Robert Bradshaw, Craig Chambers, Slava Chernyak, Rafael Fernández-Moctezuma, Reuven Lax, Sam McVeety, Daniel Mills, Frances Perry, Eric Schmidt 0001, Sam Whittle |
Proc. VLDB Endow. | 6 |
| 2013 | MillWheel: Fault-Tolerant Stream Processing at Internet ScaleabstractMillWheel is a framework for building low-latency data-processing applications that is widely used at Google. Users specify a directed computation graph and application code for individual nodes, and the system manages persistent state and the continuous flow of records, all within the envelope of the framework's fault-tolerance guarantees. This paper describes MillWheel's programming model as well as its implementation. The case study of a continuous anomaly detector in use at Google serves to motivate how many of MillWheel's features are used. MillWheel's programming model provides a notion of logical time, making it simple to write time-based aggregations. MillWheel was designed from the outset with fault tolerance and scalability in mind. In practice, we find that MillWheel's unique combination of scalability, fault tolerance, and a versatile programming model lends itself to a wide variety of problems at Google. Tyler Akidau, Alex Balikov, Kaya Bekiroglu, Slava Chernyak, Josh Haberman, Reuven Lax, Sam McVeety, Daniel Mills, Paul Nordstrom, Sam Whittle |
Proc. VLDB Endow. | 6 |