Jack Ng

dblp:41/4574 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 55% Cloud and datacenter computing · 22% Storage systems · 22%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 43% Transaction processing and concurrency control · 43% Query processing and optimization · 14%

Topics — the 6 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transaction processing and concurrency control › concurrency control
locking protocols
0.712023
Taurus MM: bringing multi-master to the cloud · Proc. VLDB Endow. 2023
Distributed systems › replication › replication and fault tolerance
replication and recovery
0.412020
Taurus Database: How to be Fast, Available, and Frugal in the Cloud · SIGMOD Conference 2020
Query processing and optimization
cardinality estimation
0.112007
Progressive optimization in a shared-nothing parallel database · SIGMOD Conference 2007
Query processing and optimization › adaptive query processing
progressive optimization
0.112007
Progressive optimization in a shared-nothing parallel database · SIGMOD Conference 2007
Query processing and optimization
query optimization
0.112007
Progressive optimization in a shared-nothing parallel database · SIGMOD Conference 2007
Parallel and multicore computing › parallel computing › parallel database systems
shared-nothing parallel database machine
0.012007
Progressive optimization in a shared-nothing parallel database · SIGMOD Conference 2007

Methods — techniques the papers use, named apart from their topics

constant-time snapshots · 0.4append-only storage · 0.4parallel checkpoint operators · 0.1voting schemes · 0.1voting scheme · 0.1
YearPublicationVenuePosition
2023 Taurus MM: bringing multi-master to the cloud
abstract
A single-master database has limited update capacity because a single node handles all updates. A multi-master database potentially has higher update capacity because the load is spread across multiple nodes. However, the need to coordinate updates and ensure durability can generate high network traffic. Reducing network load is particularly important in a cloud environment where the network infrastructure is shared among thousands of tenants. In this paper, we present Taurus MM, a shared-storage multi-master database optimized for cloud environments. It implements two novel algorithms aimed at reducing network traffic plus a number of additional optimizations. The first algorithm is a new type of distributed clock that combines the small size of Lamport clocks with the effective support of distributed snapshots of vector clocks. The second algorithm is a new hybrid page and row locking protocol that significantly reduces the number of lock requests sent over the network. Experimental results on a cluster with up to eight masters demonstrate superior performance compared to Aurora multi-master and CockroachDB.
Alex Depoutovitch, Per-Åke Larson, Jack Ng, Guanzhu Xiong, Paul Lee, Emad Boctor, Samiao Ren, Lengdong Wu, Calvin Sun
Proc. VLDB Endow.4
2020 Taurus Database: How to be Fast, Available, and Frugal in the Cloud
abstract
Using cloud Database as a Service (DBaaS) offerings instead of on-premise deployments is increasingly common. Key advantages include improved availability and scalability at a lower cost than on-premise alternatives. In this paper, we describe the design of Taurus, a new multi-tenant cloud database system. Taurus separates the compute and storage layers in a similar manner to Amazon Aurora and Microsoft Socrates and provides similar benefits, such as read replica support, low network utilization, hardware sharing and scalability. However, the Taurus architecture has several unique advantages. Taurus offers novel replication and recovery algorithms providing better availability than existing approaches using the same or fewer replicas. Also, Taurus is highly optimized for performance, using no more than one network hop on critical paths and exclusively using append-only storage, delivering faster writes, reduced device wear, and constant-time snapshots. This paper describes Taurus and provides a detailed description and analysis of the storage node architecture, which has not been previously available from the published literature.
Alex Depoutovitch, Jin Chen 0006, Per-Åke Larson, Jack Ng, Wenlin Cui
SIGMOD Conference6
2007 Progressive optimization in a shared-nothing parallel database
abstract
Commercial enterprise data warehouses are typically implemented on parallel databases due to the inherent scalability and performance limitation of a serial architecture. Queries used in such large data warehouses can contain complex predicates as well as multiple joins, and the resulting query execution plans generated by the optimizer may be sub-optimal due to mis-estimates of row cardinalities. Progressive optimization (POP) is an approach to detect cardinality estimation errors by monitoring actual cardinalities at run-time and to recover by triggering re-optimization with the actual cardinalities measured. However, the original serial POP solution is based on a serial processing architecture, and the core ideas cannot be readily applied to a parallel shared-nothing environment. Extending the serial POP to a parallel environment is a challenging problem since we need to determine when and how we can trigger re-optimization based on cardinalities collected from multiple independent nodes. In this paper, we present a comprehensive and practical solution to this problem, including several novel voting schemes whether to trigger re-optimization, a mechanism to reuse local intermediate results across nodes as a partitioned materialized view, several flavors of parallel checkpoint operators, and parallel checkpoint processing methods using efficient communication protocols. This solution has been prototyped in a leading commercial parallel DBMS. We have performed extensive experiments using the TPC-H benchmark and a real-world database. Experimental results show that our solution has negligible runtime overhead and accelerates the performance of complex OLAP queries by up to a factor of 22.
Wook-Shin Han, Jack Ng, Volker Markl, Holger Kache, Mokhtar Kandil
SIGMOD Conference2