Qingchao Cai

dblp:25/5210 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-2999-1293ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 44% Memory systems · 24% Parallel and multicore computing · 14%
Databases, data mining, and information retrieval
6 papers
Data mining · 44% Indexing and storage engines · 27% Web and social media mining · 14%
Artificial intelligence
1 paper
Efficient and distributed learning · 87% Deep learning architectures and training · 13%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › statistical analysis
cohort analysis
0.932018
Effective Temporal Dependence Discovery in Time Series Data · Proc. VLDB Endow. 2018
Cohort Analysis with Ease · SIGMOD Conference 2018
Cohort Query Processing · Proc. VLDB Endow. 2016
Storage systems › storage management
versioned storage
0.822020
ForkBase: Immutable, Tamper-evident Storage Substrate for Branchable Applications · ICDE 2020
ForkBase: An Efficient Storage Engine for Blockchain and Forkable Applications · Proc. VLDB Endow. 2018
Indexing and storage engines
in-memory index
0.622018
A Comprehensive Performance Evaluation of Modern In-Memory Indices · ICDE 2018
Parallelizing Skip Lists for In-Memory Multi-Core Database Systems · ICDE 2017
Storage systems › data reduction
data deduplication
0.412020
ForkBase: Immutable, Tamper-evident Storage Substrate for Branchable Applications · ICDE 2020
Web and social media mining
user behavior analysis
0.312018
Cohort Analysis with Ease · SIGMOD Conference 2018
Performance modeling and evaluation
benchmarking
0.312018
A Comprehensive Performance Evaluation of Modern In-Memory Indices · ICDE 2018
Memory systems › cache coherence
cache coherence protocol
0.312018
Efficient Distributed Memory Management with RDMA and Caching · Proc. VLDB Endow. 2018
Memory systems › shared memory › distributed shared memory
distributed memory management
0.312018
Efficient Distributed Memory Management with RDMA and Caching · Proc. VLDB Endow. 2018
Memory systems › shared memory
distributed shared memory
0.312018
Efficient Distributed Memory Management with RDMA and Caching · Proc. VLDB Endow. 2018
Distributed systems
fault tolerance
0.312018
Efficient Distributed Memory Management with RDMA and Caching · Proc. VLDB Endow. 2018
Storage systems
logging
0.312018
Efficient Distributed Memory Management with RDMA and Caching · Proc. VLDB Endow. 2018
Storage systems
storage engine
0.312018
ForkBase: An Efficient Storage Engine for Blockchain and Forkable Applications · Proc. VLDB Endow. 2018
Parallel and multicore computing
concurrent data structures
0.312017
Parallelizing Skip Lists for In-Memory Multi-Core Database Systems · ICDE 2017
Parallel and multicore computing › parallel algorithms
parallel indexing
0.312017
Parallelizing Skip Lists for In-Memory Multi-Core Database Systems · ICDE 2017
Machine learning › Efficient and distributed learning
distributed training
0.212015
SINGA: A Distributed Deep Learning Platform · ACM Multimedia 2015
Machine learning › Efficient and distributed learning › distributed training › parallelization
model partitioning
0.212015
SINGA: A Distributed Deep Learning Platform · ACM Multimedia 2015
Data mining
temporal analysis
0.112018
Effective Temporal Dependence Discovery in Time Series Data · Proc. VLDB Endow. 2018
Blockchain and cryptocurrency security › blockchain data management
blockchain storage
0.112018
ForkBase: An Efficient Storage Engine for Blockchain and Forkable Applications · Proc. VLDB Endow. 2018
Data models and query languages › SQL
SQL extension
0.112016
Cohort Query Processing · Proc. VLDB Endow. 2016

Methods — techniques the papers use, named apart from their topics

breadth-first search · 0.6SIMD · 0.6hash-based deduplication · 0.4content-addressed storage · 0.4distributed query processing · 0.3access methods · 0.3columnar evaluation · 0.2neural net partitioning · 0.2
YearPublicationVenuePosition
2020 ForkBase: Immutable, Tamper-evident Storage Substrate for Branchable Applications
abstract
Data collaboration activities typically require systematic or protocol-based coordination to be scalable. Git, an effective enabler for collaborative coding, has been attested for its success in countless projects around the world. Hence, applying the Git philosophy to general data collaboration beyond coding is motivating. We call it Git for data. However, the original Git design handles data at the file granule, which is considered too coarse-grained for many database applications. We argue that Git for data should be co-designed with database systems. To this end, we developed ForkBase to make Git for data practical. ForkBase is a distributed, immutable storage system designed for data version management and data collaborative operation. In this demonstration, we show how ForkBase can greatly facilitate collaborative data management and how its novel data deduplication technique can improve storage efficiency for archiving massive data versions.
Qian Lin 0002, Kaiyuan Yang 0003, Tien Tuan Anh Dinh, Qingchao Cai, Gang Chen 0001, Beng Chin Ooi, Pingcheng Ruan, Sheng Wang 0011, Zhongle Xie, Meihui Zhang 0001, Olafs Vandans
ICDE4
2019 MemepiC: Towards a Unified In-Memory Big Data Management System
abstract
In-memory data management systems have recently gained a lot of attraction due to cheaper and faster DRAM and other hardware advancement. However, these systems are either pure storage systems with online data query service, or just offline batch processing systems with data analytics functionality. Heavy data movement (e.g., data loading) occurs in order to analyze the data. In this paper, we propose an innovative in-memory data management system-MemepiC, which unifies both online data query and data analytics functionality, allowing low-latency storage service and efficient in-situ data analytics. We also explore the emerging RDMA technique in the context of in-memory data management systems, by designing an RDMA-based communication protocol for message delivery inside MemepiC, and proposing to overlap execution and RDMA communication. Extensive experiments are conducted to show the superior performance of MemepiC in terms of both the storage and the data analytics services, compared against other in-memory systems.
Qingchao Cai, Hao Zhang 0029, Wentian Guo, Gang Chen 0001, Beng Chin Ooi, Kian-Lee Tan, Weng-Fai Wong
IEEE Trans. Big Data1
2018 A Comprehensive Performance Evaluation of Modern In-Memory Indices
abstract
Due to poor cache utilization and latching contention, the B-tree like structures, which have been heavily used in traditional databases, are not suitable for modern in-memory databases running over multi-core infrastructure. To address the problem, several in-memory indices, such as FAST, Masstree, BwTree, ART and PSL, have recently been proposed, and they show good performance in concurrent settings. Given the various design choices and implementation techniques being adopted by these indices, it is therefore important to understand how these techniques and properties actually affect the indexing performance. To this end, we conduct a comprehensive performance study to compare these indices from multiple perspectives, including query throughput, scalability, latency, memory consumption as well as cache/branch miss rate, using various query workloads with different characteristics. Our results indicate that there is no one-size-fits-all solution. For example, PSL achieves better query throughput for most settings, but occupies more memory space and can incur a large overhead in updating the index. Nevertheless, the huge performance gain renders the exploitation of modern hardware features indispensable for modern database indices.
Zhongle Xie, Qingchao Cai, Gang Chen 0001, Rui Mao 0001, Meihui Zhang 0001
ICDE2
2018 Cohort Analysis with Ease
abstract
The tremendous volume of user behavior records generated in various domains provides data analysts new opportunities to mine valuable insights into user behavior. Cohort analysis, which aims to find user behavioral trends hidden in time series, is one of the most commonly used techniques. Since traditional database systems suffer from both operability and efficiency when processing cohort analysis queries, we proposed COHANA, a query processing system specialized for cohort analysis. In order to make COHANA easy-to-use, we present a comprehensive and powerful tool in this demo, covering the major use cases in cohort analysis with intuitive and accessible operations. Analysts can easily adapt COHANA to their own use with provided visualizations which can help verify their analysis assumptions and inconspicuous trends hidden in user behavior data.
Zhongle Xie, Qingchao Cai, Gene Yan Ooi, Weilong Huang, Beng Chin Ooi
SIGMOD Conference2
2018 Efficient Distributed Memory Management with RDMA and Caching
abstract
Recent advancements in high-performance networking interconnect significantly narrow the performance gap between intra-node and inter-node communications, and open up opportunities for distributed memory platforms to enforce cache coherency among distributed nodes. To this end, we propose GAM, an efficient distributed in-memory platform that provides a directory-based cache coherence protocol over remote direct memory access (RDMA). GAM manages the free memory distributed among multiple nodes to provide a unified memory model, and supports a set of user-friendly APIs for memory operations. To remove writes from critical execution paths, GAM allows a write to be reordered with the following reads and writes, and hence enforces partial store order (PSO) memory consistency. A light-weight logging scheme is designed to provide fault tolerance in GAM. We further build a transaction engine and a distributed hash table (DHT) atop GAM to show the ease-of-use and applicability of the provided APIs. Finally, we conduct an extensive micro benchmark to evaluate the read/write/lock performance of GAM under various workloads, and a macro benchmark against the transaction engine and DHT. The results show the superior performance of GAM over existing distributed memory platforms.
Qingchao Cai, Wentian Guo, Hao Zhang 0029, Divyakant Agrawal, Gang Chen 0001, Beng Chin Ooi, Kian-Lee Tan, Yong Meng Teo, Sheng Wang 0011
Proc. VLDB Endow.1
2018 Effective Temporal Dependence Discovery in Time Series Data
abstract
To analyze user behavior over time, it is useful to group users into cohorts, giving rise to cohort analysis. We identify several crucial limitations of current cohort analysis, motivated by the unmet need for temporal dependence discovery. To address these limitations, we propose a generalization that we call recurrent cohort analysis. We introduce a set of operators for recurrent cohort analysis and design access methods specific to these operators in both single-node and distributed environments. Through extensive experiments, we show that recurrent cohort analysis when implemented using the proposed access methods is up to six orders faster than one implemented as a layer on top of a database in a single-node setting, and two orders faster than one implemented using Spark SQL in a distributed setting.
Qingchao Cai, Zhongle Xie, Gang Chen 0001, H. V. Jagadish, Beng Chin Ooi, Meihui Zhang 0001
Proc. VLDB Endow.1
2018 ForkBase: An Efficient Storage Engine for Blockchain and Forkable Applications
abstract
Existing data storage systems offer a wide range of functionalities to accommodate an equally diverse range of applications. However, new classes of applications have emerged, e.g., blockchain and collaborative analytics, featuring data versioning, fork semantics, tamper-evidence or any combination thereof. They present new opportunities for storage systems to efficiently support such applications by embedding the above requirements into the storage. In this paper, we present ForkBase , a storage engine designed for blockchain and forkable applications. By integrating core application properties into the storage, ForkBase not only delivers high performance but also reduces development effort. The storage manages multiversion data and supports two variants of fork semantics which enable different fork worklflows. ForkBase is fast and space efficient, due to a novel index class that supports efficient queries as well as effective detection of duplicate content across data objects, branches and versions. We demonstrate ForkBase 's performance using three applications: a blockchain platform, a wiki engine and a collaborative analytics application. We conduct extensive experimental evaluation against respective state-of-the-art solutions. The results show that ForkBase achieves superior performance while significantly lowering the development effort.
Sheng Wang 0011, Tien Tuan Anh Dinh, Qian Lin 0002, Zhongle Xie, Meihui Zhang 0001, Qingchao Cai, Gang Chen 0001, Beng Chin Ooi, Pingcheng Ruan
Proc. VLDB Endow.6
2017 Parallelizing Skip Lists for In-Memory Multi-Core Database Systems
abstract
Due to the coarse granularity of data accesses and the heavy use of latches, indices in the B-tree family are not efficient for in-memory databases, especially in the context of today's multi-core architecture. In this paper, we study the parallelizability of skip lists for the parallel and concurrent environment, and present PSL, a Parallel in-memory Skip List that lends itself naturally to the multi-core environment, particularly with non-uniform memory access. For each query, PSL traverses the index in a Breadth-First-Search (BFS) to find the list node with the matching key, and exploits SIMD processing to speed up this process. Furthermore, PSL distributes incoming queries among multiple execution threads disjointly and uniformly to eliminate the use of latches and achieve a high parallelizability. The experimental results show that PSL is comparable to a readonly index, FAST, in terms of read performance, and outperforms ART and Masstree respectively by up to 30% and 5x for a variety of workloads.
Zhongle Xie, Qingchao Cai, H. V. Jagadish, Beng Chin Ooi, Weng-Fai Wong
ICDE2
2016 Cohort Query Processing
abstract
Modern Internet applications often produce a large volume of user activity records. Data analysts are interested in cohort analysis, or finding unusual user behavioral trends, in these large tables of activity records. In a traditional database system, cohort analysis queries are both painful to specify and expensive to evaluate. We propose to extend database systems to support cohort analysis. We do so by extending SQL with three new operators. We devise three different evaluation schemes for cohort query processing. Two of them adopt a non-intrusive approach. The third approach employs a columnar based evaluation scheme with optimizations specifically designed for cohort query processing. Our experimental results confirm the performance benefits of our proposed columnar database system, compared against the two non-intrusive approaches that implement cohort queries on top of regular relational databases.
Dawei Jiang, Qingchao Cai, Gang Chen 0001, H. V. Jagadish, Beng Chin Ooi, Kian-Lee Tan, Anthony K. H. Tung
Proc. VLDB Endow.2
2015 SINGA: A Distributed Deep Learning Platform
abstract
Deep learning has shown outstanding performance in various machine learning tasks. However, the deep complex model structure and massive training data make it expensive to train. In this paper, we present a distributed deep learning system, called SINGA, for training big models over large datasets. An intuitive programming model based on the layer abstraction is provided, which supports a variety of popular deep learning models. SINGA architecture supports both synchronous and asynchronous training frameworks. Hybrid training frameworks can also be customized to achieve good scalability. SINGA provides different neural net partitioning schemes for training large models. SINGA is an Apache Incubator project released under Apache License 2.
Beng Chin Ooi, Kian-Lee Tan, Sheng Wang 0011, Wei Wang 0059, Qingchao Cai, Gang Chen 0001, Jinyang Gao, Zhaojing Luo, Anthony K. H. Tung, Yuan Wang 0003, Zhongle Xie, Meihui Zhang 0001, Kaiping Zheng
ACM Multimedia5
2014 Virt Cache: Managing Virtual Disk Performance Variation in Distributed File Systems for the Cloud
abstract
As Applications are moved from physical servers to virtual machines sharing storage resources, they experience large variation in I/O latencies. While maintaining average performance in such virtualized environments is important to conform to service level agreements (SLA), cloud users also expect their applications to have minimum variation in tail end latencies like 90th percentile latency for predictable performance. This becomes a challenging problem as the deviation in the application's 90th percentile I/O latency from average latency under storage resource sharing (VM consolidation) can be very high. We show through experiments under VM consolidation that during peak loads this latency variation from average can be as much as 5 times compared to when the application has exclusive access to the storage devices. This variation in performance exists for both Hard drives (HDD) and Solid state drives (SSD). To minimize this large latency variation, we propose a dynamic I/O redirection and caching mechanism called Virt Cache. Virt Cache can pro-actively detect storage device contention at the storage server and temporarily redirect the peaking virtual disk workload to a dynamically instantiated distributed read-write cache. We have implemented our system in Gluster FS, a commonly used distributed file system deployed as a backing store in the cloud. Our system can achieve from 50% to 83% reduction in the 90th percentile latency deviation from average compared to previous work as we move from low load conditions to peak non uniform consolidated VM workloads. With our Virt Cache system, Cloud providers can guarantee predictable performance for the cloud users as if their application has exclusive access to the storage resources.
Rajesh Vellore Arumugam, Quanqing Xu, Haixiang Shi, Qingchao Cai, Yonggang Wen 0001
CloudCom4