Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jason Gustafson

dblp:157/9687 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 49% Cloud and datacenter computing · 44% Storage systems · 8%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.822023
Kora: A Cloud-Native Event Streaming Platform for Kafka · Proc. VLDB Endow. 2023
Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka · SIGMOD Conference 2021
Cloud and datacenter computing › cloud platform
cloud-native platform
0.712023
Kora: A Cloud-Native Event Streaming Platform for Kafka · Proc. VLDB Endow. 2023
Data stream processing › fault tolerance
exactly-once semantics
0.512021
Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka · SIGMOD Conference 2021
Cloud and datacenter computing
cluster resource management and scheduling
0.212023
Kora: A Cloud-Native Event Streaming Platform for Kafka · Proc. VLDB Endow. 2023
Distributed systems › consistency models
consistency guarantees
0.112021
Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka · SIGMOD Conference 2021
Storage systems › file systems › write-optimized file system
log-structured file system
0.112021
Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka · SIGMOD Conference 2021

Methods — techniques the papers use, named apart from their topics

transactional write protocol · 1.0idempotent write protocol · 1.0distributed systems design · 0.7
YearPublicationVenuePosition
2023 Kora: A Cloud-Native Event Streaming Platform for Kafka
abstract
Event streaming is an increasingly critical infrastructure service used in many industries and there is growing demand for cloud-native solutions. Confluent Cloud provides a massive scale event streaming platform built on top of Apache Kafka with tens of thousands of clusters running in 70+ regions across AWS, Google Cloud, and Azure. This paper introduces Kora , the cloud-native platform for Apache Kafka at the core of Confluent Cloud. We describe Kora's design that enables it to meet its cloud-native goals, such as reliability, elasticity, and cost efficiency. We discuss Kora's abstractions which allow users to think in terms of their workload requirements and not the underlying infrastructure, and we discuss how Kora is designed to provide consistent, predictable performance across cloud environments with diverse capabilities.
Anna Povzner, Prince Mahajan, Jason Gustafson, Jun Rao, Ismael Juma, Feng Min, Shriram Sridharan, Nikhil Bhatia, Gopi K. Attaluri, Adithya Chandra, Stanislav Kozlovski, Rajini Sivaram, Lucas Bradstreet, Bob Barrett, Dhruvil Shah, David Jacot, David Arthur, Manveer Chawla, Ron Dagostino, Colin Mccabe, Manikumar Reddy Obili, Kowshik Prakasam, Jose Garcia Sancio, Alok Nikhil
Proc. VLDB Endow.3
2021 Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka
abstract
An increasingly important system requirement for distributed stream processing applications is to provide strong correctness guarantees under unexpected failures and out-of-order data so that its results can be authoritative (not needing complementary batch results). Although existing systems have put a lot of effort into addressing some specific issues, such as consistency and completeness, how to enable users to make flexible and transparent trade-off decisions among correctness, performance, and cost still remains a practical challenge. Specifically, similar mechanisms are usually applied to tackle both consistency and completeness, which can result in unnecessary performance penalties. We present Apache Kafka's core design for stream processing, which relies on its persistent log architecture as the storage and inter-processor communication layers to achieve correctness guarantees. Kafka Streams, a scalable stream processing client library in Apache Kafka, defines the processing logic as read-process-write cycles in which all processing state updates and result outputs are captured as log appends. Idempotent and transactional write protocols are utilized to guarantee exactly-once semantics. Furthermore, revision-based speculative processing is employed to emit results as soon as possible while handling out-of-order data. We also demonstrate how Kafka Streams behaves in practice with large-scale deployments and performance insights exhibiting its flexible and low-overhead trade-offs.
Guozhang Wang, Ayusman Dikshit, Jason Gustafson, Matthias Sax, John Roesler, Sophie Blee-Goldman, Bruno Cadonna, Apurva Mehta, Varun Madan, Jun Rao
SIGMOD Conference4