Jun Yang 0022

dblp:181/2799-22 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Storage systems · 43% Memory systems · 37% Distributed systems · 20%
Databases, data mining, and information retrieval
1 paper
Data mining · 77% Database system architecture and tuning · 23%

Topics — the 27 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
2.272021
Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent Memory · Proc. VLDB Endow. 2021
Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Storage systems
crash consistency
0.732018
Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018
Transactional NVM cache with high performance and crash consistency · SC 2017
Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017
Storage systems
file systems
0.732018
Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017
Transactional NVM cache with high performance and crash consistency · SC 2017
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Distributed systems › fault tolerance
checkpointing
0.712023
OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023
Distributed systems › distributed machine learning
parameter server
0.712023
OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023
Memory systems › non-volatile memory
persistent memory
0.712023
OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023
Distributed systems › fault tolerance › checkpointing
synchronous checkpointing
0.712023
OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023
Storage systems › file systems
journaling file system
0.622017
Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017
Transactional NVM cache with high performance and crash consistency · SC 2017
Data mining › dimensionality reduction
feature extraction
0.512021
Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent Memory · Proc. VLDB Endow. 2021
Memory systems › non-volatile memory › persistent memory
persistent memory indexing
0.512021
Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent Memory · Proc. VLDB Endow. 2021
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory
0.422017
Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM · ACM Trans. Storage 2017
Transactional NVM cache with high performance and crash consistency · SC 2017
Storage systems › data reduction
data deduplication
0.312018
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Storage systems › data reduction › data deduplication
inline deduplication
0.312018
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Storage systems › file systems
versioning
0.312018
Persisting RB-Tree into NVM in a Consistency Perspective · ACM Trans. Storage 2018
Storage systems › file systems › file system design
persistent memory file system
0.322018
NV-Tree: Reducing Consistency Cost for NVM-based Single Level Systems · FAST 2015
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Storage systems › indexing
b+-tree
0.212016
NV-Tree: A Consistent and Workload-Adaptive Tree Structure for Non-Volatile Memory · IEEE Trans. Computers 2016
Distributed systems
data consistency
0.212016
NV-Tree: A Consistent and Workload-Adaptive Tree Structure for Non-Volatile Memory · IEEE Trans. Computers 2016
Storage systems
key-value storage
0.212016
NV-Tree: A Consistent and Workload-Adaptive Tree Structure for Non-Volatile Memory · IEEE Trans. Computers 2016
Storage systems › flash and SSD › flash memory management › flash translation layer
address mapping
0.212015
Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive · ACM Trans. Storage 2015
Storage systems
flash and SSD
0.212015
Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive · ACM Trans. Storage 2015
Storage systems › flash and SSD › flash memory management
flash translation layer
0.212015
Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive · ACM Trans. Storage 2015
Storage systems › flash and SSD › flash memory management
garbage collection
0.212015
Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive · ACM Trans. Storage 2015
Memory systems › memory management
single level store
0.212015
NV-Tree: Reducing Consistency Cost for NVM-based Single Level Systems · FAST 2015
Database system architecture and tuning › main-memory database
distributed in-memory database
0.112021
Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent Memory · Proc. VLDB Endow. 2021
Storage systems
storage reliability
0.112018
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Memory systems › non-volatile memory › write reliability
write endurance
0.112018
NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory · IEEE Trans. Computers 2018
Performance modeling and evaluation › workload characterization
workload classification
0.112015
Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive · ACM Trans. Storage 2015

Methods — techniques the papers use, named apart from their topics

persistent skiplist · 1.0distributed query processing · 1.0pipeline processing · 0.7checkpointing · 0.7DRAM cache · 0.7workload-adaptive fingerprinting · 0.3transactional metadata consistency · 0.3cascade-versioning · 0.3lightweight transaction scheme · 0.3journaling · 0.3
YearPublicationVenuePosition
2023 OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory
abstract
In this paper, we present OpenEmbedding, a distributed parameter server system for deep learning recommendation models (DLRM) workloads. In order to support rapid growth in the number of features and the model size (Terabytes are common) of DLRM workloads, OpenEmbedding takes advantage of emerging persistent memory (PMem) to address scalability and reliability issues in training DLRMs. Compared to DRAM, PMem can have much lower per-GB cost, higher density, and non-volatility, while with slightly low access performance to DRAM. OpenEmbedding uses DRAM as cache and PMem as storage for the sparse features and develops a simple but effective pipeline processing approach to optimize the access latency of the sparse features in PMem. For reliability, we develop a lightweight synchronous checkpointing scheme that is specially co-designed with the pipelined cache to reduce the run-time overhead of checkpointing. Our evaluations on a real-world industry workload consisting of billions of parameters demonstrate 1) the effectiveness of our PMem-aware optimizations, 2) checkpointing mechanism with near-zero run-time overhead to the training performance and 3) fast recovery with up to 3.97× speedup compared to the state-of-the-art. OpenEmbedding has been deployed in hundreds of scenarios in industry within 4Paradigm, and is open-sourced1.
Cheng Chen 0008, Jun Yang 0022, Mian Lu, Zhao Zheng, Bingsheng He, Weng-Fai Wong, Liang You, Penghao Sun, Yuping Zhao, Fenghua Hu, Andy Rudoff
ICDE3
2021 Optimizing An In-memory Database System For AI-powered On-line Decision Augmentation Using Persistent Memory
abstract
On-line decision augmentation (OLDA) has been considered as a promising paradigm for real-time decision making powered by Artificial Intelligence (AI). OLDA has been widely used in many applications such as real-time fraud detection, personalized recommendation, etc. On-line inference puts real-time features extracted from multiple time windows through a pre-trained model to evaluate new data to support decision making. Feature extraction is usually the most time-consuming operation in many OLDA data pipelines. In this work, we started by studying how existing in-memory databases can be leveraged to efficiently support such real-time feature extractions. However, we found that existing in-memory databases cost hundreds or even thousands of milliseconds. This is unacceptable for OLDA applications with strict real-time constraints. We therefore propose FEDB ( F eature E ngineering D ata b ase), a distributed in-memory database system designed to efficiently support on-line feature extraction. Our experimental results show that FEDB can be one to two orders of magnitude faster than the state-of-the-art in-memory databases on real-time feature extraction. Furthermore, we explore the use of the Intel Optane DC Persistent Memory Module (PMEM) to make FEDB more cost-effective. When comparing the proposed PMEM-optimized persistent skiplist to the FEDB using DRAM+SSD, PMEM-based FEDB can shorten the tail latency up to 19.7%, reduce the recovery time up to 99.7%, and save up to 58.4% total cost of a real OLDA pipeline.
Cheng Chen 0008, Jun Yang 0022, Mian Lu, Taize Wang, Zhao Zheng, Yuqiang Chen, Wenyuan Dai, Bingsheng He, Weng-Fai Wong, Guoan Wu, Yuping Zhao, Andy Rudoff
Proc. VLDB Endow.2
2018 NV-Dedup: High-Performance Inline Deduplication for Non-Volatile Memory
abstract
The byte-addressable non-volatile memory (NVM) is a promising medium for data storage. NVM-oriented file systems have been designed to explore NVM's performance potential. Meanwhile, applications may write considerable duplicate data. For NVM, a removal of duplicate data can promote space efficiency, improve write endurance, and potentially improve the performance by avoidance of repeatedly writing the same data. However, we have observed severe performance degradations when implementing a state-of-the-art inline deduplication algorithm in an NVM-oriented file system. A quantitative analysis reveals that, with NVM, 1) the conventional way to manage deduplication metadata for block devices, particularly in light of consistency, is inefficient, and, 2) the performance with deduplication becomes more subject to fingerprint calculations. We hence propose a deduplication algorithm called NV-Dedup. NV-Dedup manages deduplication metadata in a fine-grained, CPU and NVM-favored way, and preserves the metadata consistency with a lightweight transactional scheme. It also does workload-adaptive fingerprinting based on an analytical model and a transition scheme among fingerprinting methods to reduce calculation penalties. We have built a prototype of NV-Dedup in the Persistent Memory File System (PMFS). Experiments show that, NV-Dedup not only substantially saves NVM space, but also boosts the performance of PMFS by up to 2.1x.
Chundong Wang 0001, Qingsong Wei, Jun Yang 0022, Cheng Chen 0008, Yechao Yang, Mingdi Xue
IEEE Trans. Computers3
2018 Persisting RB-Tree into NVM in a Consistency Perspective
abstract
Byte-addressable non-volatile memory (NVM) is going to reshape conventional computer systems. With advantages of low latency, byte-addressability, and non-volatility, NVM can be directly put on the memory bus to replace DRAM. As a result, both system and application softwares have to be adjusted to perceive the fact that the persistent layer moves up to the memory. However, most of the current in-memory data structures will be problematic with consistency issues if not well tuned with NVM. This article places emphasis on an important in-memory structure that is widely used in computer systems, i.e., the Red/Black-tree (RB-tree). Since it has a long and complicated update process, the RB-tree is prone to inconsistency problems with NVM. This article presents an NVM-compatible consistent RB-tree with a new technique named cascade-versioning . The proposed RB-tree (i) is all-time consistent and scalable and (ii) needs no recovery procedure after system crashes. Experiment results show that the RB-tree for NVM not only achieves the aim of consistency with insignificant spatial overhead but also yields comparable performance to an ordinary volatile RB-tree.
Chundong Wang 0001, Qingsong Wei, Lingkun Wu, Sibo Wang 0001, Cheng Chen 0008, Xiaokui Xiao, Jun Yang 0022, Mingdi Xue, Yechao Yang
ACM Trans. Storage7
2017 Transactional NVM cache with high performance and crash consistency
abstract
The byte-addressable non-volatile memory (NVM) is new promising storage medium. Compared to NAND flash memory, the next-generation NVM not only preserves the durability of stored data but has much shorter access latencies. An architect can utilize the fast and persistent NVM as an external disk cache. Regarding the system's crash consistency, a prevalent journaling file system needs to run atop an NVM disk cache. However, the performance is severely impaired by redundant efforts in achieving crash consistency in both file system and disk cache. Therefore, we propose a new mechanism called transactional NVM disk cache (Tinca). In brief, Tinca jointly guarantees consistency of file system and disk cache and removes the performance penalty of file system journaling with a lightweight transaction scheme. Evaluations confirm that Tinca significantly outperforms state-of-the-art design by up to 2.5X in local and cluster tests without causing any inconsistency issue.
Qingsong Wei, Chundong Wang 0001, Cheng Chen 0008, Yechao Yang, Jun Yang 0022, Mingdi Xue
SC5
2017 Optimizing File Systems with Fine-grained Metadata Journaling on Byte-addressable NVM
abstract
Journaling file systems have been widely adopted to support applications that demand data consistency. However, we observed that the overhead of journaling can cause up to 48.2% performance drop under certain kinds of workloads. On the other hand, the emerging high-performance, byte-addressable Non-volatile Memory (NVM) has the potential to minimize such overhead by being used as the journal device. The traditional journaling mechanism based on block devices is nevertheless unsuitable for NVM due to the write amplification of metadata journal we observed. In this article, we propose a fine-grained metadata journal mechanism to fully utilize the low-latency byte-addressable NVM so that the overhead of journaling can be significantly reduced. Based on the observation that conventional block-based metadata journal contains up to 90% clean metadata that is unnecessary to be journalled, we design a fine-grained journal format for byte-addressable NVM which contains only modified metadata. Moreover, we redesign the process of transaction committing, checkpointing, and recovery in journaling file systems utilizing the new journal format. Therefore, thanks to the reduced amount of ordered writes for journals, the overhead of journaling can be reduced without compromising the file system consistency. To evaluate our fine-grained metadata journaling mechanism, we have implemented a journaling file system prototype based on Ext4 and JBD2 in Linux. Experimental results show that our NVM-based fine-grained metadata journaling is up to 15.8 × faster than the traditional approach under FileBench workloads.
Cheng Chen 0008, Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Mingdi Xue
ACM Trans. Storage2
2016 Extending SSD Lifetime with Persistent In-Memory Metadata Management
abstract
Flash-based solid state drive (SSD) is now widely deployed to speed up data intensive applications. However, I/O amplifications caused by file system metadata and journaling shorten the lifetime of SSD. In this paper, a mechanism named Persistent In-memory Metadata Management (referred to as PIMM) is proposed to reduce I/O traffics to SSD by exploiting the persistency and byte-addressability of Non-volatile Memory (NVM). The PIMM decouples data and metadata access paths, putting data on SSD and metadata in NVM at runtime. Thus, metadata is accessed in byte-addressable manner via the memory bus and metadata I/O is eliminated because metadata in NVM is not flushed back to SSD anymore. The PIMM is prototyped on real NVDIMM platform. Extensive evaluations on implemented prototype show that the proposed PIMM reduces the block erase for SSD by up to 91% and improves performance for different workloads.
Qingsong Wei, Cheng Chen 0008, Mingdi Xue, Chundong Wang 0001, Jun Yang 0022
CLUSTER5
2016 Fine-grained metadata journaling on NVM
abstract
Journaling file systems have been widely used where data consistency must be assured. However, we observed that the overhead of journaling can cause up to 48.2% performance drop under certain kinds of workloads. On the other hand, the emerging high-performance, byte-addressable Non-volatile Memory (NVM) has the potential to minimize such overhead by being used as the journal device. The traditional journaling mechanism based on block devices is nevertheless unsuitable for NVM due to the write amplification of metadata journal we observed. In this paper, we propose a fine-grained metadata journal mechanism to fully utilize the low-latency byte-addressable NVM so that the overhead of journaling can be significantly reduced. Based on the observation that conventional block-based metadata journal contains up to 90% clean metadata that is unnecessary to be journalled, we design a fine-grained journal format for byte-addressable NVM which contains only modified metadata. Moreover, we redesign the process of transaction committing, checkpointing and recovery in journaling file systems utilizing the new journal format. Therefore, thanks to the reduced amount of ordered writes to NVM, the overhead of journaling can be reduced without compromising the file system consistency. Experimental results show that our NVM-based fine-grained metadata journaling is up to 15.8× faster than the traditional approach under FileBench workloads.
Cheng Chen 0008, Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Mingdi Xue
MSST2
2016 NV-Tree: A Consistent and Workload-Adaptive Tree Structure for Non-Volatile Memory
abstract
The non-volatile memory (NVM) which can provide DRAM-like performance and disk-like persistency has the potential to build single-level systems by replacing both DRAM and disk. Keeping data consistency in such systems is non-trivial because memory writes may be reordered by CPU. Although ordered memory writes for achieving data consistency can be implemented using the memory fence and the CPU cache line flush instructions, they introduce a significant overhead (more than 10X slower in performance). In this paper, we focus on an important and common data structure, B$^+$Tree. Based on our quantitative analysis for consistent tree structures, we propose NV-Tree, a consistent, cache-optimized and workload-adaptive B$^+$Tree variant with significantly reduced consistency cost (up to 96 percent reduction in CPU cache line flush). To further optimize NV-Tree under various workloads, we propose a workload-adaptive scheme in which the sizes of individual nodes can be dynamically adjusted to improve the performance over time. We implement and evaluate NV-Tree and NV-Store, a key-value store based on NV-Tree, on an NVDIMM server. NV-Tree outperforms the state-of-art consistent tree structures by up to 12X under write-intensive workloads. NV-Store increases the throughput by up to 7.3X under YCSB workloads compared to Redis.
Jun Yang 0022, Qingsong Wei, Chundong Wang 0001, Cheng Chen 0008, Khai Leong Yong, Bingsheng He
IEEE Trans. Computers1
2015 NV-Tree: Reducing Consistency Cost for NVM-based Single Level Systems
Jun Yang 0022, Qingsong Wei, Cheng Chen 0008, Chundong Wang 0001, Khai Leong Yong, Bingsheng He
FAST1
2015 Accelerating Cloud Storage System with Byte-Addressable Non-Volatile Memory
abstract
As building block for cloud storage, distributed file system uses underlying local file systems to manage objects. However, the underlying file system, which is limited by metadata and journaling I/O, significantly affects the performance of the distributed file system. This paper presents an NVM-based file system (referred to as NV-Booster) to accelerate object access for storage node. The NV-Booster leverages byte-addressability and persistency of nonvolatile memory (NVM) to speedup metadata accesses and file system journaling. With NV-Booster, metadata is kept in NVM and accessed in byte-addressable manner through memory bus, while object is stored on hard disk and accessed from I/O bus. In addition, proposed NV-Booster enables fast object search and mapping between object ID and on-disk location with an efficient in-memory namespace management. NV-Booster is implemented in kernel space with NVDIMM and has been extensively evaluated under various workloads. Our experiments show that NV-Booster improves Ceph performance up to 10X, compared to the Ceph with existing local file systems.
Qingsong Wei, Mingdi Xue, Jun Yang 0022, Chundong Wang 0001, Cheng Chen 0008
ICPADS3
2015 How to be consistent with persistent memory? An evaluation approach
abstract
The advent of the byte-addressable, non-volatile memory (NVM) has initiated the design of new data management strategies to utilize it as the persistent memory (PM). One way to manage the PM is via an in-memory file system. The consistency of the in-memory file system may nevertheless be compromised from directly exposing the PM to the CPU, because data are likely to be flushed from the CPU cache to the PM in an order that is different from the order in which they have been programed to be. As a result, in spite of classic consistency mechanisms, such as journaling and Copy-on-Write, file systems for the PM have to seek support of cacheline flush and memory fence instructions, e.g., clflush and sfence, to achieve ordered writes. On the other hand, manipulating the PM as a consistent block device with conventional file systems is also doable. The pros and cons of two approaches, however, have not been thoroughly investigated yet. We hence do so with extensive evaluations and detailed analyses. Our aim of this paper is to inspire how the PM shall be managed, especially from the performance perspective.
Chundong Wang 0001, Qingsong Wei, Jun Yang 0022, Cheng Chen 0008, Mingdi Xue
NAS3
2015 Z-MAP: A Zone-Based Flash Translation Layer with Workload Classification for Solid-State Drive
abstract
Existing space management and address mapping schemes for flash-based Solid-State-Drive (SSD) operate either at page or block granularity, with inevitable limitations in terms of memory requirement, performance, garbage collection, and scalability. To overcome these limitations, we proposed a novel space management and address mapping scheme for flash referred to as Z-MAP, which manages flash space at granularity of Zone. Each Zone consists of multiple numbers of flash blocks. Leveraging workload classification, Z-MAP explores Page-mapping Zone (Page Zone) to store random data and handle a large number of partial updates, and Block-mapping Zone (Block Zone) to store sequential data and lower the overall mapping table. Zones are dynamically allocated and a mapping scheme for a Zone is determined only when it is allocated. Z-MAP uses a small part of Flash memory or phase change memory as a streaming Buffer Zone to log data sequentially and migrate data into Page Zone or Block Zone based on workload classification. A two-level address mapping is designed to reduce the overall mapping table and address translation latency. Z-MAP classifies data before it is permanently stored into Flash memory so that different workloads can be isolated and garbage collection overhead can be minimized. Z-MAP has been extensively evaluated by trace-driven simulation and a prototype implementation on OpenSSD. Our benchmark results conclusively demonstrate that Z-MAP can achieve up to 76% performance improvement, 81% mapping table reduction, and 88% garbage collection overhead reduction compared to existing Flash Translation Layer (FTL) schemes.
Qingsong Wei, Cheng Chen 0008, Mingdi Xue, Jun Yang 0022
ACM Trans. Storage4
2014 CBM: A cooperative buffer management for SSD
abstract
Random writes significantly limit the application of Solid State Drive (SSD) in the I/O intensive applications such as scientific computing, Web services, and database. While several buffer management algorithms are proposed to reduce random writes, their ability to deal with workloads mixed with sequential and random accesses is limited. In this paper, we propose a cooperative buffer management scheme referred to as CBM, which coordinates write buffer and read cache to fully exploit temporal and spatial localities among I/O intensive workload. To improve both buffer hit rate and destage sequentiality, CBM divides write buffer space into Page Region and Block Region. Randomly written data is put in the Page Region at page granularity, while sequentially written data is stored in the Block Region at block granularity. CBM leverages threshold-based migration to dynamically classify random write from sequential writes. When a block is evicted from write buffer, CBM merges the dirty pages in write buffer and the clean pages in read cache belonging to the evicted block to maximize the possibility of forming full block write. CBM has been extensively evaluated with simulation and real implementation on OpenSSD. Our testing results conclusively demonstrate that CBM can achieve up to 84% performance improvement and 85% garbage collection overhead reduction compared to existing buffer management schemes.
Qingsong Wei, Cheng Chen 0008, Jun Yang 0022
MSST3