Jinghuan Yu

dblp:229/3006 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0001-9729-1841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 STEM: Streaming-Based FPGA Acceleration for Large-Scale Compactions in LSM KV
abstract
Log-Structured-Merge-tree (LSM-tree) has been extensively adopted because of its exceptional write efficiency and high space utilization. Compaction is invoked periodically in LSM-tree based key-value(LSM KV) systems to maintain good system performance. As the size of LSM-KV grows, large-scale compaction is now frequently seen. Compaction throughput significantly degrades with larger inputs, leading to frequent write stalls and decrement in overall write throughput. This paper proposes STEM, a stream-based compaction framework with FPGA to address this issue. A clean-cut algorithm is introduced to enable streaming-based compaction for large-scale data. With a multi-unit pipeline and dynamic pipeline schedule, STEM can handle large-scale compaction tasks efficiently. Based on the experiment result, the compaction throughput of STEM can achieve$27\times$on average and up to$35\times$improvement compared with the current RocksDB compaction,$2.09\times$to$2.27\times$improvement compared with the state-of-the-art FPGA accelerator.
Dongdong Tang, Weilan Wang, Yu Mao 0001, Jinghuan Yu, Tei-Wei Kuo, Chun Jason Xue
ICDE4
2023 ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved Performance
Jinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason Xue
FAST1
2022 NIC-QF: A design of FPGA based Network Interface Card with Query Filter for big data systems
Jinyu Zhan, Wei Jiang 0016, Ying Li 0130, Junting Wu, Jianping Zhu 0003, Jinghuan Yu
Future Gener. Comput. Syst.6
2022 NFL: Robust Learned Index via Distribution Transformation
abstract
Recent works on learned index open a new direction for the indexing field. The key insight of the learned index is to approximate the mapping between keys and positions with piece-wise linear functions. Such methods require partitioning key space for a better approximation. Although lots of heuristics are proposed to improve the approximation quality, the bottleneck is that the segmentation overheads could hinder the overall performance. This paper tackles the approximation problem by applying a distribution transformation to the keys before constructing the learned index. A two-stage Normalizing-Flow-based Learned index framework (NFL) is proposed, which first transforms the original complex key distribution into a near-uniform distribution, then builds a learned index leveraging the transformed keys. For effective distribution transformation, we propose a Numerical Normalizing Flow (Numerical NF). Based on the characteristics of the transformed keys, we propose a robust After-Flow Learned Index (AFLI). To validate the performance, comprehensive evaluations are conducted on both synthetic and real-world workloads, which shows that the proposed NFL produces the highest throughput and the lowest tail latency compared to the state-of-the-art learned indexes.
Shangyu Wu, Yufei Cui, Jinghuan Yu, Xuan Sun 0003, Tei-Wei Kuo, Chun Jason Xue
Proc. VLDB Endow.3
2022 Accelerating Queries of Big Data Systems by Storage-Side CPU-FPGA Co-Design
abstract
As a promising technology of big data systems, storage and computing separated architecture has attracted increasing attention of famous companies, such as Tencent, IBM, Facebook, and Microsoft. Under this new architecture, conventional query engines like Hive and Presto choose all the original data from storage nodes and send them to computing nodes to be filtered, causing high data transmission overhead and great I/O bandwidth fluctuation. To address this problem, we design a novel data processing framework to prefilter data on storage side, and then propose a CPU-FPGA (field-programmable gate array) co-design to accelerate the queries with the purpose of reducing the communication overheads and the workloads of computing nodes. To obtain the optimal efficiency of CPU-FPGA co-processing, a workload-aware task scheduler is designed to allocate query tasks to CPU or FPGA according to the estimation of the filtering data size and processing time of query tasks. A data projection scheme is designed to support data in RCFile format which is widely used in modern systems, such as Tencent and Facebook applications. To make full use of the high parallelism of FPGA, we formulate the SQL conditions of combined predicates into Boolean parameters, and design two filtering schemes on FPGA (i.e., parallel sequential filter for fix-length data type and parallel pipeline filter for variable-length data type). Experiments on the TPC-H benchmark and Tencent data set demonstrate the efficiency of our approach, which can save up to 72.28% and 80.16% of time overheads compared with Presto and Hive, respectively.
Jinyu Zhan, Wei Jiang 0016, Ying Li 0130, Junting Wu, Jianping Zhu 0003, Jinghuan Yu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Accelerating data filtering for database using FPGA
Xuan Sun 0003, Chun Jason Xue, Jinghuan Yu, Tei-Wei Kuo, Xue (Steve) Liu
J. Syst. Archit.3
2020 Position: Synergetic effects of Software and Hardware Parameters on the LSM system
Jinghuan Yu, Heejin Yoon, Sam H. Noh, Young-ri Choi, Chun Jason Xue
HotStorage1
2020 FPGA-based Compaction Engine for Accelerating LSM-tree Key-Value Stores
abstract
With the rapid growth of big data, LSM-tree based key-value stores are widely applied due to its high efficiency in write performance. Compaction plays a critical role in LSM-tree, which merges old data and could significantly reduce the overall throughput of the whole system especially for write-intensive workloads. Hardware acceleration for database is a popular trend in recent years. In this paper, we design and implement an FPGA-based compaction engine to accelerate compaction in LSM-tree based key-value stores. To take full advantage of the pipeline mechanism on FPGA, the key-value separation and index-data block separation strategies are proposed. In order to improve the compaction performance, the bandwidth of FPGA-chip is fully utilized. In addition, the proposed acceleration engine is integrated with a classic LSM-tree based store without modifications on the original storage format. The experimental results demonstrate that the proposed FPGA-based compaction engine can achieve up to 92.0x acceleration ratio compared with CPU baseline, and achieve up to 6.4x improvement on the throughput of random writes.
Xuan Sun 0003, Jinghuan Yu, Zimeng Zhou, Chun Jason Xue
ICDE2
2020 Branch-aware data variable allocation for energy optimization of hybrid SRAM+NVM SPM☆
Jinyu Zhan, Wei Jiang 0016, Jiayu Yu, Jinghuan Yu
J. Syst. Archit.5
2018 Persistence improvement for distributed cache with NVM based storage system: work-in-progress
Wei Jiang 0016, Jinyu Zhan, Jinghuan Yu, Liugen Xu
CASES4