Yikang Xu

dblp:37/3729 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 69% Distributed systems · 18% Cloud and datacenter computing · 12%
Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
storage reliability
0.522019
Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures · USENIX ATC 2019
Erasure Coding in Windows Azure Storage · USENIX ATC 2012
Indexing and storage engines
key-value store
0.512021
ArkDB: A Key-Value Engine for Scalable Cloud Storage Services · SIGMOD Conference 2021
Indexing and storage engines
write amplification reduction
0.512021
ArkDB: A Key-Value Engine for Scalable Cloud Storage Services · SIGMOD Conference 2021
Storage systems › distributed storage
disaggregated storage
0.512021
ArkDB: A Key-Value Engine for Scalable Cloud Storage Services · SIGMOD Conference 2021
Storage systems › file systems
distributed file system
0.512021
ArkDB: A Key-Value Engine for Scalable Cloud Storage Services · SIGMOD Conference 2021
Cloud and datacenter computing
cloud storage
0.432021
ArkDB: A Key-Value Engine for Scalable Cloud Storage Services · SIGMOD Conference 2021
Erasure Coding in Windows Azure Storage · USENIX ATC 2012
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Storage systems › storage reliability
failure characterization
0.412019
Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures · USENIX ATC 2019
Storage systems
flash and SSD
0.412019
Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures · USENIX ATC 2019
Storage systems › storage reliability
erasure coding
0.112012
Erasure Coding in Windows Azure Storage · USENIX ATC 2012
Distributed systems
fault tolerance
0.112011
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Distributed systems › replication
geo-replication
0.112011
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Distributed systems › fault tolerance
high availability
0.112011
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Distributed systems
replication
0.112011
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Distributed systems › consistency models
strong consistency
0.112011
Windows Azure Storage: a highly available cloud storage service with strong consistency · SOSP 2011
Electronic design automation › hardware verification and test
debugging
0.011997
Replaying Distributed Programs without Message Logging · HPDC 1997
Distributed systems › fault tolerance
checkpointing
0.011997
Replaying Distributed Programs without Message Logging · HPDC 1997

Methods — techniques the papers use, named apart from their topics

partition split and merge · 1.0page mapping table · 1.0garbage collection · 1.0field data analysis · 0.4failure log analysis · 0.4
YearPublicationVenuePosition
2021 ArkDB: A Key-Value Engine for Scalable Cloud Storage Services
abstract
Persistent key-value stores play a crucial role in enabling internet-scale services. At Alibaba Cloud, scale-out cloud storage services including Object Storage Service, File Storage Service and Tablestore are built on distributed key-value stores. Key challenges in the design of the underlying key-value engine for these services lie in utilization of disaggregated storage, supporting write and range query-heavy workloads, and balancing of scalability, availability and resource usage. This paper presents ArkDB, a key-value engine designed to address these challenges by combining advantages of both LSM tree and Bw-tree, and leveraging advances in hardware technologies. Built on top of Pangu, an append-only distributed file system, ArkDB's innovations include shrinkable page mapping table, clear separation of system and user states for fast recovery, write amplification reduction, efficient garbage collection and lightweight partition split and merge. Experimental results demonstrate ArkDB's improvements over existing designs. Compared with Bw-tree, ArkDB efficiently stabilizes the mapping table size despite continuous write working set growth. Compared with RocksDB, an LSM tree-based key-value engine, ArkDB increases ingestion throughput by 2.16x, while reducing write amplification by 3.1x. It outperforms RocksDB by 52% and 37% respectively on a write-heavy workload and a range query-intensive workload of the Yahoo! Cloud Serving Benchmark. Experiments running in Tablestore in a cluster environment further demonstrate ArkDB's performance on Pangu and its efficient partition split/merge support.
Zhu Pang, Qingda Lu, Rui Wang 0002, Yikang Xu, Jiesheng Wu
SIGMOD Conference5
2019 Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures
Erci Xu, Mai Zheng, Yikang Xu, Jiesheng Wu
USENIX ATC4
2012 Erasure Coding in Windows Azure Storage
Cheng Huang 0002, Huseyin Simitci, Yikang Xu, Aaron Ogus, Brad Calder, Parikshit Gopalan, Jin Li 0001, Sergey Yekhanin
USENIX ATC3
2011 Windows Azure Storage: a highly available cloud storage service with strong consistency
abstract
Windows Azure Storage (WAS) is a cloud storage system that provides customers the ability to store seemingly limitless amounts of data for any duration of time. WAS customers have access to their data from anywhere at any time and only pay for what they use and store. In WAS, data is stored durably using both local and geographic replication to facilitate disaster recovery. Currently, WAS storage comes in the form of Blobs (files), Tables (structured storage), and Queues (message delivery). In this paper, we describe the WAS architecture, global namespace, and data model, as well as its resource provisioning, load balancing, and replication systems.
Brad Calder, Ju Wang 0010, Aaron Ogus, Niranjan Nilakantan, Arild Skjolsvold, Sam McKelvie, Yikang Xu, Shashwat Srivastav, Jiesheng Wu, Huseyin Simitci, Jaidev Haridas, Chakravarthy Uddaraju, Hemal Khatri, Andrew Edwards, Vaman Bedekar, Shane Mainali, Rafay Abbasi, Mian Fahim ul Haq, Muhammad Ikram ul Haq, Deepali Bhardwaj, Sowmya Dayanand, Anitha Adusumilli, Marvin McNett, Sriram Sankaran, Kavitha Manivannan, Leonidas Rigas
SOSP7
1997 Replaying Distributed Programs without Message Logging
abstract
Debugging long program runs can be difficult because of the delays required to repeatedly re-run the execution. Even a moderately long run of five minutes can incur aggravating delays. To address this problem, techniques exist that allow re-executing a distributed program from intermediate points by using combinations of checkpointing and message logging. In this paper we explore another idea: how to support replay without logging the contents of any message. When no messages are logged, the set of global states from which replay is possible is constrained, and it has been unknown how to compute this set without exhaustively searching the space of all global states, whose size is exponential in the number of processes. We present a simple and efficient hybrid on-the-fly/post-mortem algorithm for detecting the necessary and sufficient conditions under which parts of the execution can be replayed without message logs. A small amount of trace (two vectors) is recorded at each checkpoint and a fast post-mortem algorithm computes global states from which replay can begin. This algorithm is independent of the checkpointing technique used.
Robert H. B. Netzer, Yikang Xu
HPDC2