Xuan Sun 0003

dblp:117/5602-3 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-6023-7400ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Pipette: Efficient Fine-Grained Reads for SSDs
abstract
Big data applications, such as recommendation system and social network, often generate a huge number of fine-grained reads to the storage. Block-oriented storage devices upon the traditional storage system rely on the paging mechanism to migrate pages to the host DRAM, tending to suffer from these fine-grained read operations in terms of I/O traffic as well as performance. Motivated by this challenge, an efficient fine-grained read framework, Pipette, is proposed in this article as an extension to the traditional I/O framework. With adaptive design for caching, merging, and scheduling, Pipette explores locality and acceleration for fine-grained read requests to establish an efficient byte-granular read path upon the dedicated byte-addressable interface. When the Pipette prototype on an SSD runs popular workloads, we measured throughput gains by up to 50% and 54% with traffic reduction in the range of$41.3\times $and$56.5\times $.
Shuhan Bai, Hu Wan 0001, Yun Huang 0005, Xuan Sun 0003, Fei Wu 0005, Changsheng Xie 0001, Hung-Chih Hsieh, Tei-Wei Kuo, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Pipette: efficient fine-grained reads for SSDs
abstract
Big data applications, such as recommendation system and social network, often generate a huge number of fine-grained reads to the storage. Block-oriented storage devices tend to suffer from these fine-grained read operations in terms of I/O traffic as well as performance. Motivated by this challenge, a fine-grained read framework, Pipette, is proposed in this paper, as an extension to the traditional I/O framework. With an adaptive caching design, Pipette framework offers a tremendous reduction in I/O traffic as well as achieves significant performance gain. A Pipette prototype was implemented with Ext4 file system on an SSD for two real-world applications, where the I/O throughput is improved by 31.6% and 33.5%, and the I/O traffic is reduced by 95.6% and 93.6%, respectively.
Shuhan Bai, Hu Wan 0001, Yun Huang 0005, Xuan Sun 0003, Fei Wu 0005, Changsheng Xie 0001, Hung-Chih Hsieh, Tei-Wei Kuo, Chun Jason Xue
DAC4
2022 $p$LPAQ: Accelerating LPAQ Compression on FPGA
abstract
In recent years, the demand for data storage space has increased dramatically due to the exponential growth of data volume. Data compression is of great significance since it saves data storage space and reduces data transfer demand. Compression algorithms based on statistical models have a much higher compression ratio than dictionary-based methods, but the high computational time cost of statistical modeling limits their wider application. In this paper, we introduce pLPAQ, an FPGA-based design of a powerful compression algorithm LPAQ based on statistical models. A novel hardware accelerator is proposed to speed up LPAQ by fully utilizing the parallelism of FPGA. Experimental results show that the proposed design can achieve a throughput of 12 MB/s on Xilinx Virtex Plus UltraScale XCVU9P card, 25x faster than executing on AMD Ryzen R7 4800U at 2.8 GHz and 80x faster compared with the naive FPGA implementation on average.
Dongdong Tang, Xuan Sun 0003, Nan Guan, Tei-Wei Kuo, Chun Jason Xue
FPT2
2022 RM-SSD: In-Storage Computing for Large-Scale Recommendation Inference
abstract
To meet the strict service level agreement requirements of recommendation systems, the entire set of embeddings in recommendation systems needs to be loaded into the memory. However, as the model and dataset for production-scale recommendation systems scale up, the size of the embeddings is approaching the limit of memory capacity. Limited physical memory constrains the algorithms that can be trained and deployed, posing a severe challenge for deploying advanced recommendation systems. Recent studies offload the embedding lookups into SSDs, which targets the embedding-dominated recommendation models. This paper takes it one step further and proposes to offload the entire recommendation system into SSD with in-storage computing capability. The proposed SSD-side FPGA solution leverages a low-end FPGA to speed up both the embedding-dominated and MLP-dominated models with high resource efficiency. We evaluate the performance of the proposed solution with a prototype SSD. Results show that we can achieve 20-100× throughput improvement compared with the baseline SSD and 1.5-15× improvement compared with the state-of-art.
Xuan Sun 0003, Hu Wan 0001, Qiao Li 0001, Chia-Lin Yang, Tei-Wei Kuo, Chun Jason Xue
HPCA1
2022 NFL: Robust Learned Index via Distribution Transformation
abstract
Recent works on learned index open a new direction for the indexing field. The key insight of the learned index is to approximate the mapping between keys and positions with piece-wise linear functions. Such methods require partitioning key space for a better approximation. Although lots of heuristics are proposed to improve the approximation quality, the bottleneck is that the segmentation overheads could hinder the overall performance. This paper tackles the approximation problem by applying a distribution transformation to the keys before constructing the learned index. A two-stage Normalizing-Flow-based Learned index framework (NFL) is proposed, which first transforms the original complex key distribution into a near-uniform distribution, then builds a learned index leveraging the transformed keys. For effective distribution transformation, we propose a Numerical Normalizing Flow (Numerical NF). Based on the characteristics of the transformed keys, we propose a robust After-Flow Learned Index (AFLI). To validate the performance, comprehensive evaluations are conducted on both synthetic and real-world workloads, which shows that the proposed NFL produces the highest throughput and the lowest tail latency compared to the state-of-the-art learned indexes.
Shangyu Wu, Yufei Cui, Jinghuan Yu, Xuan Sun 0003, Tei-Wei Kuo, Chun Jason Xue
Proc. VLDB Endow.4
2021 Accelerating data filtering for database using FPGA
Xuan Sun 0003, Chun Jason Xue, Jinghuan Yu, Tei-Wei Kuo, Xue (Steve) Liu
J. Syst. Archit.1
2020 FPGA-based Compaction Engine for Accelerating LSM-tree Key-Value Stores
abstract
With the rapid growth of big data, LSM-tree based key-value stores are widely applied due to its high efficiency in write performance. Compaction plays a critical role in LSM-tree, which merges old data and could significantly reduce the overall throughput of the whole system especially for write-intensive workloads. Hardware acceleration for database is a popular trend in recent years. In this paper, we design and implement an FPGA-based compaction engine to accelerate compaction in LSM-tree based key-value stores. To take full advantage of the pipeline mechanism on FPGA, the key-value separation and index-data block separation strategies are proposed. In order to improve the compaction performance, the bandwidth of FPGA-chip is fully utilized. In addition, the proposed acceleration engine is integrated with a classic LSM-tree based store without modifications on the original storage format. The experimental results demonstrate that the proposed FPGA-based compaction engine can achieve up to 92.0x acceleration ratio compared with CPU baseline, and achieve up to 6.4x improvement on the throughput of random writes.
Xuan Sun 0003, Jinghuan Yu, Zimeng Zhou, Chun Jason Xue
ICDE1