VLDB 2026 Research / reviewers in the wild / expert
John Macmillan
dblp:224/6790
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
stream processing systems |
0.3 | 1 | 2018 | Challenges and Experiences in Building an Efficient Apache Beam Runner For IBM Streams · Proc. VLDB Endow. 2018 |
Cloud and datacenter computing › big data platform
stream processing engine |
0.3 | 1 | 2018 | Challenges and Experiences in Building an Efficient Apache Beam Runner For IBM Streams · Proc. VLDB Endow. 2018 |
Methods — techniques the papers use, named apart from their topics
state indexing · 0.7garbage collection · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Challenges and Experiences in Building an Efficient Apache Beam Runner For IBM StreamsabstractThis paper describes the challenges and experiences in the development of IBM Streams runner for Apache Beam. Apache Beam is emerging as a common stream programming interface for multiple computing engines. Each participating engine implements a runner to translate Beam applications into engine-specific programs. Hence, applications written with the Beam SDK can be executed on different underlying stream computing engines, with negligible migration penalty. IBM Streams is a widely-used enterprise streaming platform. It has a rich set of connectors and toolkits for easy integration of streaming applications with other enterprise applications. It also supports a broad range of programming language interfaces, including Java, C++, Python, Stream Processing Language (SPL) and Apache Beam. This paper focuses on our solutions to efficiently support the Beam programming abstractions in IBM Streams runner. Beam organizes data into discrete event time windows. This design, on the one hand, supports out-of-order data arrivals, but on the other hand, forces runners to maintain more states, which leads to higher space and computation overhead. IBM Streams runner mitigates this problem by efficiently indexing inter-dependent states, garbage-collecting stale keys, and enforcing bundle sizes. We also share performance concerns in Beam that could potentially impact applications. Evaluations show that IBM Streams runner outperforms Flink runner and Spark runner in most scenarios when running the Beam NEXMark benchmarks. IBM Streams runner is available for download from IBM Cloud Streaming Analytics service console. Paul Gerver, John Macmillan, Daniel Debrunner, William Marshall, Kun-Lung Wu |
Proc. VLDB Endow. | 3 |