VLDB 2026 Research / reviewers in the wild / expert
Hui Sun 0002
dblp:31/3925-2
· DBLP profile ↗
57ranked-venue papers
47as first author
46since 2021 · last 2026
0000-0003-1811-1318ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 32 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Learned Data Compression via Dual-Stream Feature DecouplingabstractHuidong Ma, Xinyan Shi, Sun Hui, Xiaofei Yue, Xiaoguang Liu, Gang Wang, Wentong Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Huidong Ma, Xinyan Shi, Hui Sun 0002, Xiaofei Yue, Xiaoguang Liu 0001, Gang Wang 0001, Wentong Cai 0001 |
ACL (1) | 3 |
| 2026 | Group-TopK: Optimizing Distributed Training on Edge Devices via Communication Compression
Yifan Wang 0005, Xiaohui Peng 0002, Haohao Ma, Hui Sun 0002, Deke Guo, Boyu Diao |
HPDC | 4 |
| 2026 | iGC: Reinforcement learning-guided intelligent garbage collection strategy for multi-tenant SSDs
Donghua Li, Hui Sun 0002, Xiao Qin 0001 |
Knowl. Based Syst. | 2 |
| 2026 | GDH+: A GPGPU-Empowered Gradient Data Hierarchy and Key-Value Separation for Optimizing LSM-Tree-Based KV StoresabstractThe rapid growth of unstructured data has driven the widespread adoption of LSM-tree-based key-value stores (KV stores). The write amplification resulting from compaction in LSM-trees causes a performance bottleneck. Existing solutions attempt to address this issue through key-value separation strategies. However, these studies fail to optimize the memory components of LSM-trees or provide efficient garbage collection (GC) strategies that achieve high performance while minimizing CPU overhead. These limitations motivate us to propose a GPGPU-empowered gradient data hierarchy and key-value separation for optimizing KV stores, named GDH+ . We utilize GPGPU acceleration for sorting and flushing operations, optimizing the memory components of the LSM-tree. Additionally, we enhance read performance with an LRU-based memory component that distinguishes between hot and cold data, in combination with an adaptive migration strategy. Furthermore, we propose an in-place GC strategy to reduce CPU overhead while maintaining high performance. GDH+ achieves a 2×, 1.5×, and 40% improvement in write performance, read performance, and CPU utilization, respectively, compared to state-of-the-art KV stores. Hui Sun 0002, Xiangxiang Jiang, Jinfeng Xu 0004, Enhui Wang, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2026 | AGC: An Adaptive Workload Burst-Aware Garbage Collection Mechanism for High-Performance SSDsabstractIn NAND flash-based solid-state drives (SSDs), system performance and quality of service tend to be degraded by frequent conflicts between garbage collection (GC) I/Os and user I/Os. To address this challenge, we propose an adaptive workload burst-Aware Garbage Collection (AGC) mechanism to build highperformance SSDs. AGC minimizes conflicts between user I/Os and GC I/Os by jointly considering access patterns in workloads, internal GC, and allocation strategies. The running time of a workload is divided into multiple five-millisecond time windows. AGC incorporates a burst I/O access pattern prediction mechanism, which anticipates I/Os patterns in subsequent time windows according to historical access patterns. Pattern predictions allow AGC to proactively buffer valid pages during GC to prevent interference with forthcoming user I/O requests. According to the status of the underlying flash memory, we implement a hybrid allocation strategy to reduce GC-induced user writes blocking while preserving read parallelism performance. AGC classifies data into four hotness levels (hot, warm, cool, and cold) according to the update interval of the same request. Then, we design a hotness-aware victim block selection policy that prioritizes hot blocks to avert valid data migration during GC, thereby reducing write amplification in SSDs. We evaluate AGC through extensive experiments driven by real-world traces. The findings confirm that compared with state-of-the-art schemes (Baseline, CachedGC, FFT-GC, HIPA, and RDA), AGC reduces read and write response time by up to 81.57% and 69.66%, respectively, with average reductions of 54.34% and 39.21%. Furthermore, AGC decreases GC counts by up to 24.31% with an average of 12.04%, thereby reducing GC overhead and extending SSD lifetime across diverse workloads. Hui Sun 0002, Haisheng Ding, Haoqiang Tong, Honggang Chai, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2026 | eCache: A Sample-Inference-Based Intelligent Cache Scheme for High-Performance SSDsabstractDRAM-based cache is a practical approach to enhancing the performance of large-capacity SSDs. Due to DRAM’s limited capacity, cache sizes are significantly smaller than the data scale of workloads – and cache replacement schemes determine the cache hit ratio and SSD performance when a cache reaches full capacity. Prior work embarked on leveraging machine learning models (ML) to predict future access patterns in workloads, aiding cache replacement decisions for intelligent cache schemes inside SSDs. Unfortunately, existing ML-based cache schemes often overlook the impacts of data granularity – including page, request, and coarse granularities – of training datasets, model inference time overhead, and computational overhead on SSD performance. To address these challenges, we are motivated to propose an sample-inference-based intelligentcachescheme – eCache. eCache accurately predicts the future reuse distance of each sampled requested pages, which can enhance the accuracy of decision-making for the cache replacement. In particular, we design a parallel framework to curb the time overhead caused by model inference. We advocate for a random-group sampling inference method that utilizes the most accurate model while reducing computational overhead. Moreover, we implement eCache on the state-of-the-art SSD simulator, MQSim, and compare it against alternative cache schemes (i.e., CCache, NCache, LAC, and VS-batch). The experimental results unveil that compared with the other cache schemes, eCache significantly reduces the average response time by up to 79.68% with an average reduction of 44.08%. When compared with the page-granularity-ML-empowered cache schemes, eCache greatly curtails computational overhead by up to 85.23% with an average reduction of 66.00%. Hui Sun 0002, Yinan Fu, Yi Zhou 0009, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2026 | JMStore: Joint Optimization of Computation and Storage Balancing in Multi-NDP Key-Value Stores with Hash-Based Data DistributionabstractIt is challenging to store and process massive unstructured data for key-value storage systems that require high concurrency, high performance, and low latency. Log-Structured Merge (LSM) trees-based Key-value stores or KV stores are widely adopted for enhanced write performance. Existing KV stores are primarily deployed on a CPU-centric architecture, which necessitates moving data to the CPU from memory or storage devices for processing. This is particularly problematic during the compaction process, which involves substantial data movement and rewrite. This tradition consumes bandwidth and computational resources, leading to write amplification and impairing system performance. Near-Data Processing (NDP) devices mitigate this issue by processing data at the storage location, thereby reducing data movement cost. Recognizing that the computational power of a single NDP device is insufficient for the demands of large-scale unstructured data processing, we propose JMStore—a multi-NDP key-value store based on a hash data organization. By offloading computational tasks to multiple NDP devices, JMStore collaboratively optimizes the compaction process to address the data movement issue as well as the mismatch between large-scale in workloads data and computational power on an NDP. We design a key-value store programming model and data organization for a multi-NDP architecture, enabling the system to leverage the hardware efficiency and parallelism of multiple NDP devices to significantly optimize system performance. We also propose a strategy to balance storage and computation resources under a hash layout. Compared to the latest single NDP KV store (PStore) and the multi-NDP KV store (MStore), JMStore demonstrates significant performance improvements under DB_Bench and YCSB-C with read-write mixed workloads: the peak improvements can reach a factor of 10. Hui Sun 0002, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Storage | 1 |
| 2025 | Genomics Data Lossless Compression with (S, K)-Mer Encoding and Deep Neural NetworksabstractLearning-based compression shows competitive compression ratios for genomics data. It often includes three types of compressors: static, adaptive and semi-adaptive. However, these existing compressors suffer from inferior compression ratios or throughput, and adaptive compressors also faces model cold-start problems. To address these issues, we propose DeepGeCo, a novel genomics data lossless adaptive compression framework with (s,k)-mer encoding and deep neural networks, involving three compression modes (MINI for static, PLUS for adaptive, ULTRA for semi-adaptive) for flexible requirements of compression ratios or throughput. In DeepGeCo, (1) we develop BiGRU and Transformer as the backbone to build Warm-Start and Supporter models in terms of cold-start problems. (2) We introduce (s,k)-mer encoding to pre-process genomics data before feeding it into the DNN model for improve model throughput, and we propose a new metric - Ranking of Throughput and Compression Ratio (RTCR) for effective encoding parameters selection. (3) We design a threshold controller and a probabilistic mixer within the backbone to balance compression ratios and model throughput. Experiments on 10 real-world datasets show that DeepGeCo's three compression modes improve up to a 22.949X average throughput and up to a 31.095% average compression ratio improvement while occupying low CPU or GPU memory. Hui Sun 0002, Liping Yi, Huidong Ma, Yongxia Sun, Yingfeng Zheng, Wenwen Cui, Meng Yan 0008, Gang Wang 0001, Xiaoguang Liu 0001 |
AAAI | 1 |
| 2025 | Multi-source Data Lossless Compression via Parallel Expansion Mapping and xLSTMabstractExplosive growth of multi-source data (MSD) poses challenges in data transmitting and storing. Neural Network (NN)-based lossless compressors are an important type of compression approaches to alleviate these problems. However, existing NN-based lossless compressors suffer from poor compression ratio and high time cost at the same time. To address these issues, we propose a novel MSD Lossless Compressor (MSDLC) with two compression stages: 1) We propose a Parallel Expansion Mapper (PEM) to map redundant pieces in MSD into unused alphabet values, which not only compresses MSD but also saves time for the next stage’s NN-based lossless compression. 2) With the mapped MSD as input, we design a NN-based lossless compressor to further improve compression ratio, where we introduce the state-of-the-art xLSTM model and design a Deep Spatial Gating Module (DSGM) as the backbone of NN. We compare MSDLC with 11 baselines on 6 real-world datasets and the results validate that MSDLC obtains the best average compression ratio and time cost. Compared with baselines, compression ratios are improved by 1.103%~113.897%, and the time costs are improved by 41.367%~73.891%. The codes can be available at https://github.com/mhuidong/MSDLC. Huidong Ma, Hui Sun 0002, Liping Yi, Xiaoguang Liu 0001, Gang Wang 0001 |
ICASSP | 2 |
| 2025 | Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTMabstractLearning-based lossless compressors have been validated to have competitive advantages in genomics data (GD) compression. However, learning-based GD-dedicated compressors typically need to be pre-trained on multi-source data and then are directly used to compress another target data, we denote them as static compressors, and they often face two challenges: limited compression ratios and bad-performed generalization due to data distribution variations. To solve these problems, we propose AGDLC, a novel Adaptive Genomics Data Lossless Compressor. It includes two critical designs: 1) We design a multiple (s, k)-mer mixer for extracting GD redundancy from multiple dimensions to improve compression ratios. 2) We introduce a recently popular XLSTM model as the backbone, which adaptively compresses GD while updating parameters, without pre-training, improving compression ratios and compression generalization at the same time. We compare AGDLC with 13 baselines on 7 real-world datasets, and the experimental results demonstrate that it achieves the best compression ratio with an average improvement of 2.162%-69.436%. The codes can be found at https://github.com/dingyanfeng/AGDLC. Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Haonan Xie, Gang Wang 0001, Xiaoguang Liu 0001 |
ICASSP | 1 |
| 2025 | PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics DatabaseabstractLearning-based lossless compressors play a crucial role in large-scale genomic database backup, storage, transmission, and management. However, their 1) inadequate compression ratio, 2) low compression & decompression throughput, and 3) poor compression robustness limit their widespread adoption and application in both industry and academia. To solve those challenges, we propose a novel Parallel Multi-Knowledge Learning-based Compressor (PMKLC) with four crucial designs: 1) We propose an automated multi-knowledge learning-based compression framework as compressors' backbone to enhance compression ratio and robustness; 2) we design a GPU-accelerated (s,k)-mer encoder to optimize compression throughput and computing resource usage; 3) we introduce data block partitioning and Step-wise Model Passing (SMP) mechanisms for parallel acceleration; 4) We design two compression modes PMKLC-S and PMKLC-M to meet the complex application scenarios, where the former runs on a resource-constrained single GPU and the latter is multi-GPU accelerated. We benchmark PMKLC-S/M and 14 baselines (7 traditional and 7 leaning-based) on 15 real-world datasets with different species and data sizes. Compared to baselines on the testing datasets, PMKLC-S/M achieve the average compression ratio improvement up to 73.609% and 73.480%, the average throughput improvement up to 3.036X and 10.710X, respectively. Besides, PMKLC-S/M also achieve the best robustness and competitive memory cost, indicating its greater stability against datasets with different probability distribution perturbations, and its strong ability to run on memory-constrained devices. Overall, PMKLC is a balanced compression solution that optimizes compression ratio, throughput, robustness, and resource consumption. PMKLC and linkages of datasets are available at https://github.com/dingyanfeng/PMKLC. Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Gang Wang 0001, Xiaoguang Liu 0001, Wentong Cai 0001 |
KDD (2) | 1 |
| 2025 | gParaKV: A GPGPU-accelerated Key-Value Separation-based KV Store with Optimized Compaction and Garbage CollectionabstractLSM-tree-based key-value stores or KV stores are widely deployed in modern cloud storage systems thanks to high data storage efficiency and retrieval capabilities. The compaction process in the LSM-tree, however, results in severe performance bottlenecks, especially in scenarios involving large volumes of data. While key-value separation methods mitigate the performance bottlenecks caused by compaction, the existing methods do not fully address merge-sorting during compaction and expensive garbage collection (GC). We propose gParaKV, a GPGPU-empowered KV store with a KV separation mechanism, leveraging the GPGPU parallel technology to accelerate merge-sorting in compaction and GC. gParaKV embraces unique features like a GPGPU bitmap structure, parallel data marking, and a parallel GC mechanism. These critical components effectively curtail the overhead of merge-sorting and GC operations by virtue of parallel computing. We compare it with state-of-the-art KV stores (e.g., RocksDB, BlobDB, Wisckey, DiffKV, UniKV, and HPDK) under various workloads. The experimental results show that gParaKV can improve the write performance and GC efficiency compared to the existing key-value separation-based KV stores. Hui Sun 0002, Xiangxiang Jiang, Xiao Qin 0001, Song Jiang 0001, Enhui Wang |
SC | 1 |
| 2025 | MSDZip: Universal Lossless Compression for Multi-source Data via Stepwise-parallel and Learning-based PredictionabstractWith the rapid development of the Internet, the huge amount of Multi-Source Data (MSD) brings challenges in data sharing and storing. Lossless data compression is the major way to solve those problems. Nowadays, neural-network technologies bring significant advantage in data modeling, making learning-based lossless compressors (LLCs) for multi-source data have emerged continuously. Compared with traditional compressors, the LLCs are more useful to catch complex redundancy patterns in MSD, and thus have great potential in enhancing compression ratio. However, existing LLCs still suffer from unsatisfactory compression ratios and lower throughput. To solve those problems, we propose a novel universal MSD lossless compressor called MSDZip via Stepwise-parallel and learning-based prediction technologies, it introduces two major designs: 1) We propose a Local-Global-Deep Mixing block in the learning-based prediction module to establish dependencies for MSD symbols, where designed Deep Mixing block solves the problem of unstable weights in the perceptual layers caused by cold-start problem to enhance the compression ratio significantly. 2) We design a Stepwise-parallel multi-GPU-accelerated compression strategy to address the compression speed and graphics memory constraints of single GPU in the face of large-scale data. The Stepwise-parallel module passes the source MSD to learning-based prediction model through the data chunking strategy, where the model of the previous chunk is used to guide the compression of the next chunk in parallel. We compare MSDZip with 5 classical learning-based and 6 traditional compressors on 12 well-studied real-world datasets. The experimental results demonstrate that MSDZip optimizes 3.418%-69.874% in terms of compression ratio and 31.171%-495.649% in terms of throughput compared to advanced LLCs. The source code of MSDZip and the linkages of the experimental datasets are available at https://github.com/mhuidong/MSDZip. Huidong Ma, Hui Sun 0002, Liping Yi, Yanfeng Ding, Xiaoguang Liu 0001, Gang Wang 0001 |
WWW | 2 |
| 2025 | A survey and benchmark evaluation for neural-network-based lossless universal compressors toward multi-source dataabstractAbstract As various types of data grow explosively, large-scale data storage, backup, and transmission become challenging, which motivates many researchers to propose efficient universal compression algorithms for multi-source data. In recent years, due to the emergence of hardware acceleration devices such as GPUs, TPUs, DPUs, and FPGAs, the performance bottleneck of neural networks (NN) has been overcome, making NN-based compression algorithms increasingly practical and popular. However, the research survey for the NN-based universal lossless compressors has not been conducted yet, and there is also a lack of unified evaluation metrics. To address the above problems, in this paper, we present a holistic survey as well as benchmark evaluations. Specifically, i) we thoroughly investigate NN-based lossless universal compression algorithms toward multi-source data and classify them into 3 types: static pre-training, adaptive, and semi-adaptive. ii) We unify 19 evaluation metrics to comprehensively assess the compression effect, resource consumption, and model performance of compressors. iii) We conduct experiments more than 4600 CPU/GPU hours to evaluate 17 state-of-the-art compressors on 28 real-world datasets across data types of text, images, videos, audio, etc. iv) We also summarize the strengths and drawbacks of NN-based lossless data compressors and discuss promising research directions. We summarize the results as the NN-based Lossless Compressors Benchmark (NNLCB, See fahaihi.github.io/NNLCB website), which will be updated and maintained continuously in the future. Hui Sun 0002, Huidong Ma, Haonan Xie, Yongxia Sun, Liping Yi, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001 |
Frontiers Comput. Sci. | 1 |
| 2025 | A+Store: An Asynchronous Parallel Compaction for Multi-NDP-Enabled Key-Value Store
Hui Sun 0002, Xiaole Liu, Yi Zhou 0009, Yinliang Yue, Xiao Qin 0001 |
J. Syst. Archit. | 1 |
| 2025 | ProckStore: An NDP-empowered key-value store with asynchronous and multi-threaded compaction scheme for optimized performance
Hui Sun 0002, Yinliang Yue, Xiao Qin 0001 |
J. Syst. Archit. | 1 |
| 2025 | HAKV: A Hotness-Aware Zone Management Approach to Optimizing Performance of LSM-tree-based Key-Value StoresabstractLog-Structured Merge tree-based key-value (KV) stores, like LevelDB and RocksDB, are extensively applied in large-scale data storage systems. This design excels in write-intensive environments by converting random writes into sequential append operations. Despite its advantages, KV stores struggle with real-world workloads where most updates in KV pairs are infrequent. The compaction process and hierarchical data organization result in high write and read amplification. To mitigate these issues, we propose HAKV – a hotness-aware zone management approach to optimizing performance of KV stores. HAKV first separates hot KV pairs from cold KV pairs, storing hot KV pairs in dedicated zones within persistent memory (PM), enabling centralized and lightweight compaction. Second, we propose a storage zone structure in PM to achieve space optimization for cold KV pairs. Third, to bolster cache hit ratio in PM, we provide a hierarchical data framework for hot KV pairs – and a recycling strategy for invalid hot KV pairs in a zone to enhance the space utilization of PM for hot KV pairs. Finally, we design a dynamic window-based adaptive adjustment mechanism for zone pool in PM to optimize the space utilization. Thus, HAKV significantly reduces write amplification while boosting overall read and write performance. The experimental results demonstrate that HAKV achieves write amplification reduction by up to 92.3%, 79.2%, 90.2%, 41.1%, 80.6%, and 62.4% compared with LevelDB, RocksDB, NoveLSM, LightKV, Wisckey, and UniKV, respectively, with average reduction rates of 89.6%, 74.4%, 84.9% 32.3%, 63.7%, and 42.5%. Furthermore, HAKV boosts random write performance by up to 54.2×, 51.5×, 44.2×, 4.3×, 3.1×, and 4.3×, respectively—and the average improvement reaches 25.8×, 20.9×, 23.9×, 2.7×, 2.5×, and 3.4×. Hui Sun 0002, Qianli Yue, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | RGKV: A GPGPU-Empowered Compaction Framework for LSM-Tree-Based KV Stores With Optimized Data Transfer and Parallel ProcessingabstractThe Log-structured merge-tree (LSM-tree), widely adopted in key-value stores (KV stores), is esteemed for its efficient write performance and superb scalability amid large-scale data processing. The compaction process of LSM-trees consumes significant computational resources, thereby becoming a bottleneck for system performance. Traditionally, compaction is handled by CPUs, but CPU processing capacity often falls short of increasing demands with the surge in data volumes. To address this challenge, existing solutions attempt to accelerate compaction using GPGPUs. Due to low GPGPU parallelism and data transfer delay in prior studies, the anticipated performance improvements have not yet been fully realized. In this paper, we bring forth RGKV – a comprehensive optimization approach to overcoming the limitations of current GPGPU-empowered KV stores. RGKV features the GPGPU-adapted contiguous memory allocation and GPGPU-optimized key-value block architecture to furnish high-efficient GPGPU parallel encoding and decoding catering to the needs of KV stores. To enhance the computational efficiency and overall performance of KV stores, RGKV employs a parallel merge-sorting algorithm to maximize the parallel processing capabilities of the GPGPU. Moreover, RGKV incorporates a data transfer module anchored on the GPUDirect storage technology – designed for KV stores – and designs an efficient data structure to substantially curtail data transfer latency between an SSD and a GPGPU, boosting data transfer speed and alleviating CPU load. The experimental results demonstrate that RGKV achieves a remarkable 4$\times$improvement in overall throughput and a 7$\times$improvement in compaction throughput compared to the state-of-the-art KV stores, while also reducing average write latency by 70.6%. Hui Sun 0002, Xiangxiang Jiang, Yinliang Yue, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2025 | iCache: An Intelligent Cache Allocation Strategy for Multitenant in High-Performance Solid-State DisksabstractThanks to high-density flash memory and high parallelism, multitenant solid-state drives (MSSDs) have become a popular high-performance storage device for enhancing cache resource utilization and reducing operational costs within these SSDs. The competition for limited cache resources inside the MSSD among multiple tenants, however, can lead to performance interference among the tenants, and prior studies focused on quality of service (QoS) in MSSDs. An efficient caching scheme is crucial for optimizing SSD performance and lifetime. Existing caching schemes aim to shorten response time by the virtue of improved cache hit rates, which offer limited performance improvement as well as low cache resource efficiency. In this article, we propose an intelligent cache allocation scheme named iCache, which employs a long short-term memory (LSTM) model to capture the I/Os access patterns of workloads and dynamically allocates cache resources inside an MSSD according to maximum benefit point (MBP) and optimal allocation point (OAP). The extensive experimental results demonstrate that iCache reduces response time by up to 87%, 24%, and 20% compared against the existing caching schemes—Shared, Justitia, and MLCache, respectively. The empirical study confirms that the new traits of iCache immensely improve system performance by enhancing the cache efficiency of MSSDs and guaranteeing fairness in performance across varying workloads. Donghua Li, Hui Sun 0002, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | MTree: A Tiering-based Key-Value Store Powered by High-performance Hierarchical Data ManagementabstractKey-value stores (KV stores) anchored on log-structured merge trees (LSM-tree) provide much improved performance under write-intensive workloads but exhibit significant write amplification (WA). The tiering compaction strategy is widely adopted in KV stores to reduce the WA. However, there are three imminent issues in tiering-based KV stores. First, excessive levels in the LSM-tree structure result in high read and write amplification. Second, without key partitioning support, these KV stores increase write tail latency. Relying solely on a hash-based key partitioning scheme is insufficient, as it does not support range queries. Furthermore, a KV store with only a key range-based partitioning scheme overlooks the distribution of keys within the key space. Third, existing KV stores with level optimization can incur high costs. A common approach to reducing the number of LSM-tree levels is to increase the MemTable size. While this reduces the number of levels, it requires more memory, leading to higher costs. Therefore, it is crucial to balance the tradeoff between level optimization and monetary expenses. To address these issues, we propose a tiering-based key-value store, leveraging high-performance hierarchical data management. The key space of our proposed KV store is partitioned into m ultiple partitions based on key ranges, with an LSM- tree in each partition. We refer to this KV store as MTree in this article. MTree reduces the number of levels in the LSM-Tree structure close to one without increasing the monetary cost. In MTree, a hierarchical structure that combines the log and LSM-tree curtails the number of levels in the LSM-tree by accumulating the data in the log without increasing the cost. With the hierarchical structure in place, we devise a long- and short-term key density distribution-aware partitioning scheme for the key range. This scheme dynamically adjusts the size of partitioning, balancing data in the partition and storing most data in large-sized logs. Then, MTree reduces the number of levels in the LSM-tree, thereby optimizing the read/write amplification. The experimental results show that MTree limits the WA to as low as two. MTree enhances the random write throughput of the state-of-the-art KV stores by up to 6.9×. Hui Sun 0002, Yinhui Chen, Yonwei Yu, Yajie Deng, Yinliang Yue, Song Jiang 0001, Xiao Qin 0001 |
ACM Trans. Storage | 1 |
| 2024 | LRCB: A Comprehensive Benchmark Evaluation of Reference-free Lossless Compression Tools for Genomics Sequencing Long Reads DataabstractThe advancement of long reads sequencing technologies has led to a significant increase in biological sequencing big data. Although several reference-free compressors are available for saving long reads data storage space, choosing the suitable one is challenging due to the shortage of thorough and systematic evaluations of their lossless compression effectiveness, both dedicated and general-purpose. In this study, we performed benchmark examinations on 30 compressors, including 11 specialized for long reads and 19 general-purpose ones, using 31 real-world datasets with differing sequencing platforms, species, and lengths. Each lossless compressor was evaluated on 13 performance measures, including compression strength, compression robustness, as well as time and peak memory required for compression and decompression. Additionally, for future long reads data compressors, we outlined investigation directions with consideration for privacy-sensitive sequences data security, hardware parallel acceleration, parameter tuning framework, and system hardware-algorithm integration design. We summarized the results as the Long Reads Compression Benchmark, available at https://github.com/fahaihi/LRCB. Hui Sun 0002, Huidong Ma, Yingfeng Zheng, Haonan Xie, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001 |
DCC | 1 |
| 2024 | A Window-Driven Compaction Mechanism in LSM-tree-based Key-Value Stores through Near-Data ProcessingabstractLSM-tree-based key-value stores or KV stores, with their advantage of sequential writes, are widely deployed as backend storage engines. LSM-trees achieve updates by merging data through compaction operations. During compaction, however, background tasks – read, write, and filter large amounts of data – compete against foreground requests for the system’s computational and I/O resources. Although numerous studies have sought to enhance system performance using high-speed storage along with computational devices, bottlenecks in CPU computation and I/Os still persist. Near-data processing (NDP) emerges as an effective solution to mitigate the bottleneck issues: the computational units of an NDP device not only expand the system’s computational resources but also significantly reduce data movement by internally performing computations. Prior collaborative efforts in the NDP device have adopted a time-aware dynamic scheduling mode. Nonetheless, there still exists a noticeable disparity in processing time between a host and its device. To address this gap, we propose WinDB – a KV store that utilizes a window-driven task allocation approach. WinDB advances task allocation in fine-grained key ranges, ensuring that the data volume of compacted SSTables matches the computing capabilities of both host and device. Apart from device-level parallelism, WinDB embraces thread-level parallelism to enhance the system’s overall performance. The experimental results unfold that WinDB bolsters the throughput by up to 4x that of RocksDB, with a significant reduction in average latency. WinDB is capable of curtailing resource contention, which in turn leads to shortened front-end response time. Hui Sun 0002, Yinliang Yue, Xiao Qin 0001 |
HPCC | 1 |
| 2024 | PQSDC: a parallel lossless compressor for quality scores data via sequences partition and run-length prediction mappingabstractMOTIVATION: The quality scores data (QSD) account for 70% in compressed FastQ files obtained from the short and long reads sequencing technologies. Designing effective compressors for QSD that counterbalance compression ratio, time cost, and memory consumption is essential in scenarios such as large-scale genomics data sharing and long-term data backup. This study presents a novel parallel lossless QSD-dedicated compression algorithm named PQSDC, which fulfills the above requirements well. PQSDC is based on two core components: a parallel sequences-partition model designed to reduce peak memory consumption and time cost during compression and decompression processes, as well as a parallel four-level run-length prediction mapping model to enhance compression ratio. Besides, the PQSDC algorithm is also designed to be highly concurrent using multicore CPU clusters. RESULTS: We evaluate PQSDC and four state-of-the-art compression algorithms on 27 real-world datasets, including 61.857 billion QSD characters and 632.908 million QSD sequences. (1) For short reads, compared to baselines, the maximum improvement of PQSDC reaches 7.06% in average compression ratio, and 8.01% in weighted average compression ratio. During compression and decompression, the maximum total time savings of PQSDC are 79.96% and 84.56%, respectively; the maximum average memory savings are 68.34% and 77.63%, respectively. (2) For long reads, the maximum improvement of PQSDC reaches 12.51% and 13.42% in average and weighted average compression ratio, respectively. The maximum total time savings during compression and decompression are 53.51% and 72.53%, respectively; the maximum average memory savings are 19.44% and 17.42%, respectively. (3) Furthermore, PQSDC ranks second in compression robustness among the tested algorithms, indicating that it is less affected by the probability distribution of the QSD collections. Overall, our work provides a promising solution for QSD parallel compression, which balances storage cost, time consumption, and memory occupation primely. AVAILABILITY AND IMPLEMENTATION: The proposed PQSDC compressor can be downloaded from https://github.com/fahaihi/PQSDC. Hui Sun 0002, Yingfeng Zheng, Haonan Xie, Huidong Ma, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001 |
Bioinform. | 1 |
| 2024 | A Machine Learning-Empowered Cache Management Scheme for High-Performance SSDsabstractNAND Flash-based solid-state drives (SSDs) have gained widespread usage in data storage thanks to their exceptional performance and low power consumption. The computational capability of SSDs has been elevated to tackle complex algorithms. Inside an SSD, a DRAM cache for frequently accessed requests reduces response time and write amplification (WA), thereby improving SSD performance and lifetime. Existing caching schemes, based on temporal locality, overlook its variations, which potentially reduces cache hit rates. Some caching schemes bolster performance via flash-aware techniques but at the expense of the cache hit rate. To address these issues, we propose a random forest machine learning Classifier-empowered Cache scheme named CCache, where I/O requests are classified into critical, intermediate, and non-critical ones according to their access status. After designing a machine learning model to predict these three types of requests, we implement a trie-level linked list to manage the cache placement and replacement. CCache safeguards critical requests for cache service to the greatest extent, while granting the highest priority to evicting request accessed by non-critical requests. CCache – considering chip state when processing non-critical requests – is implemented in an SSD simulator (SSDSim). CCache outperforms the alternative caching schemes, including LRU, CFLRU, LCR, NCache, ML_WP, and CCache_ANN, in terms of response time, WA, erase count, and hit ratio. The performance discrepancy between CCache and the OPT scheme is marginal. For example, CCache reduces the response time of the competitors by up to 41.9% with an average of 16.1%. CCache slashes erase counts by a maximum of 67.4%, with an average of 21.3%. The performance gap between CCache and and OPT is merely 2.0%-3.0%. Hui Sun 0002, Haoqiang Tong, Yinliang Yue, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2024 | LAC: A Workload Intensity-Aware Caching Scheme for High-Performance SSDsabstractInside an NAND Flash-based solid-state disk (SSD), utilizing DRAM-based write-back caching is a practical approach to bolstering the SSD performance. Existing caching schemes overlook the problem of high user I/Os intensity due to the dramatic increment of I/Os accesses. The hefty I/O intensity causes access conflict of I/O requests inside an SSD: a large number of requests are blocked to impair response time. Conventional passive update caching schemes merely replace pages upon access misses in event of full cache. Tail latency occurs facing a colossal I/O intensity. Active write-back caching schemes utilize idle time among requests coupled with free internal bandwidth to flush dirty data into flash memory in advance, lowering response time. Frequent active write-back operations, however, cause access conflict of requests – a culprit that expands write amplification (WA) and degrades SSD lifetime. We address the above issues by proposing awork Load intensity-aware and Active parallelCachingscheme - LAC - that is powered by collaborative-load awareness. LAC fends off user I/Os’ access conflict under high-I/O-intensity workloads. If the I/O intensity is low – intervals between consecutive I/O requests are large – and the target die is free, LAC actively and concurrently writes dirty data of adjacent addresses back to the die, cultivating clean data generated by the active write-back. Replacing clean data in priority can reduce response time and prevent flash transactions from being blocked. We devise a data protection method to write back cold data based on various criteria in the cache replacement and active write-backs. Thus, LAC reduces WA incurred by actively writing back hot data and extends SSD lifetime. We compare LAC against the six caching schemes (LRU, CFLRU, GCaR-LRU, MQSim, VS-Batch, and Co-Active) in the modern MQSim simulator. The results unveil that LAC trims response time and erase count by up to 78.5% and 47.8%, with an average of 64.4% and 16.6%, respectively. Hui Sun 0002, Haoqiang Tong, Yinliang Yue, Xiao Qin 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Asynchronous Compaction Acceleration Scheme for Near-data Processing-enabled LSM-tree-based KV StoresabstractLSM-tree-based key-value stores (KV stores) convert random-write requests to sequence-write ones to achieve high I/O performance. Meanwhile, compaction operations in KV stores update SSTables in forms of reorganizing low-level data components to high-level ones, thereby guaranteeing an orderly data layout in each component. Repeated writes caused by compaction (a.k.a. write amplification) impacts I/O bandwidth and overall system performance. Near-data processing (NDP) is one of the effective approaches to addressing this write-amplification issue. Most NDP-based techniques adopt synchronous parallel schemes to perform a compaction task on both the host and its NDP-enabled device. In synchronous parallel compaction schemes, the execution time of compaction is determined by a subsystem that has lower compaction performance coupled by under-utilized computing resources in a NDP framework. To solve this problem, we propose an asynchronous parallel scheme named PStore to improve the compaction performance in KV stores. In PStore, we designed a multi-tasks queue and three priority-based scheduling methods. PStore elects proper compaction tasks to be offloaded in host- and device-side compaction modules. Our proposed cross-leveled compaction mechanism mitigates write amplification induced by asynchronous compaction. PStore featured with the asynchronous compaction mechanism fully utilizes computing resources in both host- and device-side subsystems. Compared with the two popular synchronous compaction modes based on KV stores (TStore and LevelDB), our PStore immensely improves the throughput by up to a factor of 14 and 10.52 with an average of a factor of 2.09 and 1.73, respectively. Hui Sun 0002, Bendong Lou, Deyan Kong, Chaowei Zhang 0001, Jianzhong Huang 0001, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | gLSM: Using GPGPU to Accelerate Compactions in LSM-tree-based Key-value StoresabstractLog-structured-merge tree or LSM-tree is a technological underpinning in key-value (KV) stores to support a wide range of performance-critical applications. By conducting data re-organization in the background by virtue of compaction operations, the KV stores have the potential to swiftly service write requests with sequential batched disk writes and read requests for KV items constantly sorted by the compaction. Compaction demands high I/O bandwidth and CPU speed to facilitate quality service to user read/write requests. With the emergence of high-speed SSDs, CPUs are increasingly becoming a performance bottleneck. To mitigate the bottleneck limiting the KV-store’s performance and that of the applications supported by the store, we propose a system - gLSM - to leverage GPGPU to remarkably accelerate the compaction operations. gLSM fully utilizes the parallelism and computational capability inside GPGPUs to improve the compaction performance. We design a driver framework to parallelize compaction operations handled between a pair of CPU and GPGPU. We employ data independence and GPGPU-orient radix-sorting algorithm to concurrently conduct compaction. A key-value separation method is devised to slash the transfer of data volume from CPU-side memory to the GPGPU counterpart. The results reveal that gLSM improves the throughput and compaction bandwidth by up to a factor of 2.9 and 26.0, respectively, compared with the four state-of-the-art KV stores. gLSM also reduces the write latency by 73.3%. gLSM exhibits a performance improvement by up to 45% compared against its variant where there are no KV separation and collaboration sort modules. Hui Sun 0002, Jinfeng Xu 0004, Xiangxiang Jiang, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Storage | 1 |
| 2024 | TrieKV: A High-Performance Key-Value Store Design With Memory as Its First-Class CitizenabstractKey-value (KV) stores based on log-structured merge tree (LSM-tree) have been extensively studied and deployed in major information technology infrastructures. Because this type of systems is catered for KV store accessing disks, a limited disk bandwidth increases the difficulty of serving online data requests. One solution involves using a large DRAM such that frequent KV pairs are buffered and accessed from the main memory – and this solution exposes a major design drawback of the KV store: its lack of support for integrated data management in memory and on disks. For example, data in the most popular LSM-tree implementation – RocksDB – may reside in a small write buffer (MemTable) that organizes KV pairs for disk writes, a buffer cache for disk blocks, a write-ahead log on the disk for data persistence, and in various LSM levels on the disk. Without the integrated management of indexes, data, and their persistence in a hierarchical memory/disk architecture, memory is under-utilized along with missed performance optimization opportunities. We propose a KV store, TrieKV, which holistically incorporates DRAM, persistent memory (PMem), and disk with certain desired features: (1) fast in-memory access, (2) accurate identification of hot/cold data at an adaptable granularity, (3) customized memory space allocation for minimized fragmentation, (4) hotness-aware data placement across the storage hierarchy, (5) in-place data persistence in the PMem, and (6) hotness-aware LSM-tree compaction. TrieKV employs a single, integrated trie-structured index for all KV pairs in memory, where access hotness can be consistently discovered. Accordingly, the KV placement is dynamically determined according to the hotness and persistence needs of the storage hierarchy spanning the DRAM, PMem, and solid-state drive. In the experiment, we demonstrate that the 99th latency of RocksDB and NoveLSM is 38x and 6x higher than that of TrieKV, respectively. In addition, TrieKV outperforms RocksDB and NoveLSM by a factor of 5.6 and 1.7in terms of throughput, respectively. Hui Sun 0002, Deyan Kong, Song Jiang 0001, Yinliang Yue, Xiao Qin 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | EdgeBrain: A Game Based Collaborative Computational Task Offloading Framework for Edge Video AnalyticsabstractWith the development of the Internet of Everything (IoE), huge amount of data is being generated at network edge and needs to be processed in real time, such as edge video analytics applications. However, limited computing resources and energy on edge devices make it challenging to efficiently perform the intensive computing tasks. Existing solutions of offloading tasks to edge servers depend on the availability of edge servers, and they are difficult to satisfy massive computation needs. Instead, this article proposes a novel offloading framework named EdgeBrain, which offloads computational tasks from one edge device (Producer of Computation, i.e., PCO) to other edge devices (Consumers of Computation, i.e., CCOs). We formulate this offloading problem as a multi-round non-cooperative Stackelberg game, prove the existence of unique Nash Equilibrium (NE) and Stackelberg Equilibrium (SE) in the game, and design the gradient search algorithm to calculate the optimal offloading decision. The performance evaluation based on a prototype implementation of EdgeBrain in the context of a video analytics application, face detection and recognition, shows that EdgeBrain outperforms several comparable techniques in terms of execution time, frame rate, transfer data volume, and power consumption. Hui Sun 0002, Kewei Sha, Yalong Wu |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Elevating Performance of LSM-Tree-Based Key-Value Stores with Gradient Data HierarchyabstractKey-value stores are a key player of managing large-scale unstructured data in storage systems. Performance improvement of the LSM-tree structure has been extensively investigated, but current work primarily focuses on cache structural optimization rather than hot-and-cold data properties. Moreover, existing and external memory components of LSM-tree rarely have uniform hot and cold attributions. In this study, we make use of the gradient and hierarchy mechanism to optimize the components catering for cache data. We design an adaptive data migration method according to hot and cold data in the cache. We reform and expand a gradient cold-hot data hierarchy (GDH) mechanism that replaces the in-memory data structure to address the problem of missing hot and cold data attributes. The hot and cold data are placed in separate cache partitions to store hot data as far the high hierarchy as possible, reducing$\mathrm{I}/\mathrm{O}$accesses. When it comes to frequently accessed hot data, we advocate for a hotness-aware technique for data stored on a disk, where read-write performance and the cache hit rate are revamped. The experiment results reveal that our proposed GDH achieves a high cache-hit ratio and low access latency under a wide range of workloads. Hui Sun 0002, Jinfeng Xu 0004, Xiao Qin 0001 |
CLOUD | 1 |
| 2023 | SR2C: A Structurally Redundant Short Reads Collapser for Optimizing DNA Data CompressionabstractThe current redundant sequence deduplication algorithms cannot remove structural repetitive DNA short reads such as mirror, reverse, paired, and complementary palindromes in high-throughput genomics sequencing data. Moreover, these methods also cannot construct indexes to recover the original sequences, thus failing to meet the requirements of lossless compression for downstream applications. To address these problems, we propose a data structure called Cycle-Hash-Linkage (CHL) and present a CPU parallelism optimization algorithm named SR2C (Structurally Redundant Short Reads Collapser) based on CHL to improve the compression ratio of DNA sequencing data. Experimental results on actual data from the NCBI database demonstrate that SR2C achieves an average residual sequence percentage improvement of 2.556% compared to the state-of-the-art redundant sequence deduplication algorithm, Minirmd. Furthermore, SR2C cascaded optimization improves the average compression ratios of compression algorithms Pigz, PBzip2, XZ, and 7Z by 92.345%, 78.999%, 10.132%, and 7.434%, respectively. By leveraging multi-core CPU parallel computation, SR2C effectively reduces time consumption, which achieves 2-5X deduplication and recovers acceleration.The same name Linux toolkit is freely available at https://github.com/fahaihi/SR2C. Hui Sun 0002, Huidong Ma, Yingfeng Zheng, Haonan Xie, Xiaofei Wang 0001, Xiaoguang Liu 0001, Gang Wang 0001 |
ICPADS | 1 |
| 2023 | ricME: Long-Read Based Mobile Element Variant Detection Using Sequence Realignment and Identity Calculation
Huidong Ma, Hui Sun 0002, Haixiang Lin |
ISBRA | 3 |
| 2023 | PMFFRC: a large-scale genomic short reads compression optimizer via memory modeling and redundant clusteringabstractBACKGROUND: Genomic sequencing reads compressors are essential for balancing high-throughput sequencing short reads generation speed, large-scale genomic data sharing, and infrastructure storage expenditure. However, most existing short reads compressors rarely utilize big-memory systems and duplicative information between diverse sequencing files to achieve a higher compression ratio for conserving reads data storage space. RESULTS: We employ compression ratio as the optimization objective and propose a large-scale genomic sequencing short reads data compression optimizer, named PMFFRC, through novelty memory modeling and redundant reads clustering technologies. By cascading PMFFRC, in 982 GB fastq format sequencing data, with 274 GB and 3.3 billion short reads, the state-of-the-art and reference-free compressors HARC, SPRING, Mstcom, and FastqCLS achieve 77.89%, 77.56%, 73.51%, and 29.36% average maximum compression ratio gains, respectively. PMFFRC saves 39.41%, 41.62%, 40.99%, and 20.19% of storage space sizes compared with the four unoptimized compressors. CONCLUSIONS: PMFFRC rational usage big-memory of compression server, effectively saving the sequencing reads data storage space sizes, which relieves the basic storage facilities costs and community sharing transmitting overhead. Our work furnishes a novel solution for improving sequencing reads compression and saving storage space. The proposed PMFFRC algorithm is packaged in a same-name Linux toolkit, available un-limited at https://github.com/fahaihi/PMFFRC . Hui Sun 0002, Yingfeng Zheng, Haonan Xie, Huidong Ma, Xiaoguang Liu 0001, Gang Wang 0001 |
BMC Bioinform. | 1 |
| 2023 | DAC: A dynamic active and collaborative cache management scheme for solid state disks
Hui Sun 0002, Shangshang Dai, Jianzhong Huang 0001, Yinliang Yue, Xiao Qin 0001 |
J. Syst. Archit. | 1 |
| 2023 | NCache: A Machine-Learning Cache Management Scheme for Computational SSDsabstractInside a solid-state disk (SSD), cache stores frequently accessed data to shorten the user-I/O response time and reduce the number of read/write operations in flash memory, thereby improving SSD performance and lifetime. Most existing cache schemes anchor in the spatiotemporal locality of I/O requests in workloads. In the face of a long-time workload, high performance and hit rate often get lost in these caching schemes. Flash memory-aware caching schemes trade hit ratio to prolong SSD lifetime. In this article, we advocate for a machine-learning-based caching scheme named NCache to optimize both hit ratio and SSD performance. In NCache, we construct a machine learning (i.e., ML) model to predict whether data are reaccessed before being evicted from the cache. The cache replacement scheme preferentially evicts data that would not be accessed in the cache. The cache space is conserved for valid data that are likely to be repeatedly accessed. A pipelined scheme is implemented to accelerate the ML model, alleviating the time-cost of NCache. A double-linked list boosts the data addressing and cache replacement process. NCache is orthogonal to the existing caching schemes within the flash translation layer. The results validate NCache under a handful of real-world enterprise traces. Taking prn_0 as an example, NCache reduces the response time of LRU, clean first LRU (CFLRU), GCaR_LRU, GCaR_CFLRU, and LCR by up to 15% with an average of 6.4%. The erase count is slashed by 16% at the maximum. Importantly, NCache is adroit at optimizing write amplification by up to 15.9%. Hui Sun 0002, Qiao Cui, Jianzhong Huang 0001, Xiao Qin 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Improving LSM-Tree Based Key-Value Stores With Fine-Grained Compaction MechanismabstractLSM-tree-based key-value stores (KV stores) render high-performance read/write services to data-intensive applications. KV stores employ an SSTable-based Coarse-Grained Compaction (CGC) mechanism, which involves a huge amount of data that do not need to be updated, thereby bringing a high write amplification (WA) and long tail latency. To address this issue, we propose a Fine-Grained Compaction (FGC) mechanism anchored on a Log-Structured patched-Merge tree (LSpM-tree) - a new data organization that averts rewriting irrelevant data into disks amid compaction. A cluster, the basic unit in FGC, encloses several patches and a redirection table, where each patch has an array of KV regions. We devise three compaction modes powered by the LSpM-tree, and we implement a high-performance key-value store, named FGKV. The extensive experiments show that FGKV improves the random-write throughput by up to 121%, 36.8%, 38.6%, and 15.2% compared with LevelDB, RocksDB, LDC, and ALDC, respectively. FGKV lowers the WA of the alternative KV stores by up to 50%. FGKV boosts read performance by up to 122%, 51.4%, 96.6%, and 368%, respectively, and FGKV curbs the 99th percentile latency of LevelDB, RocksDB, LDC, and ALDC by up to 78.2%, 77.6%, 78.3%, and 73.1% under YCSB A, respectively. Moreover, FGKV is readily extended to the other KV stores Hui Sun 0002, Yinliang Yue, Xiao Qin 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | A Proactive On-Demand Content Placement Strategy in Edge Intelligent GatewaysabstractBandwidth-intensive applications transmit large-scale video data in the network. It causes backhaul bottlenecks and affects user experience. Deploying edge cache on an access point (AP) is a popular method to bring content files closer to end-users, but it faces significant challenges, especially in efficiently predicting and satisfying different users’ future content requests with limited cache capacity. In this article, we propose an intelligent gateway assisted edge cache deployment strategy (GACD), which jointly considers traffic usage patterns in multiple APs and the impact of new content on the cache performance. In GACD, The cache content placement problem is formulated as a many-to-one bidirectional matching problem with a dynamic quota allocation, aiming to improve cache resource utilization and minimize the average delivery latency. To address this problem, we design a heterogeneous information networks based prediction algorithm to predict end-users’ potential preference of new content files. Then, we adapt the seasonal autoregressive integrated moving average model for traffic usage prediction, and propose a many-to-one matching algorithm to achieve dynamic matching quota adjustment and efficient cache content placement. We conduct extensive real-world trace-based experiments to validate the performance of GACD. Compared with six alternative cache strategies, GACD improves the hit rate by 23.9% on average, reduces the average content delivery delay by 19.02%, and increases the accuracy by 31.02% on average. Hui Sun 0002, Kewei Sha, Shaoyuan Huang, Xiaofei Wang 0001, Weisong Shi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Poster: Ensemble Federated Edge Learning for Recommender SystemsabstractGiven the explosion of e-services, it has become critical for recommender systems (RSs) to have expected suggestions. Traditional machine learning-based recommending models provide an interface for platforms to find the most relevant items for users. Nonetheless, those models are often trained with user data from a single domain at centralized cloud, which hinders the performance of RSs, causes significant data transmission overhead, and may harm data privacy. To address these issues, in this poster, we propose an ensemble federated edge learning scheme (eFEEL) on the basis of a semi-distributed architecture design. eFEEL aims to efficiently and effectively improve RSs without breaching user data privacy. Hui Sun 0002, Kewei Sha, Yalong Wu |
SEC | 1 |
| 2022 | RT-FEND: Spark-Based Real Time FakE News DetectionabstractFake news is a rampant societal and organizational problem with various social media outlets further aggravating its spread. There is a pressing demand to assist people to identify misinformation from massive amount of news data in a timely manner. Detecting Fake news in a timely manner is critical for mitigating its impact. In this research, we propose a novel approach for detecting fake news in real time, RT-FEND (Real Time- FakE News Detection), which relies on distributed computing paradigm. The proposed methodology utilizes event and topic extraction techniques along with a topic- merging mechanism to process real time news data and reduce the number of topics for managing the curse of dimensionality. We report the findings from several experiments to compare RT-FEND with other systems to benchmark in different system settings. RT-FEND approach is more performance-improved and time-efficient in detecting fake news when compared to other fake news detection baselines. Chaowei Zhang 0001, Ashish Gupta 0004, Hui Sun 0002, Yun Li 0010, Xiao Qin 0001 |
NAS | 3 |
| 2022 | ElasticEdge: An Intelligent Elastic Edge Framework for Live Video AnalyticsabstractCloud computing and edge computing models are popularly applied in emerging applications, such as smart homes, smart parks, and connected autonomous vehicles for large-scale live video analytics. Cloud computing-based models transfer all data to the cloud for video analytics, which burdens network bandwidth and increases the data transmission overhead. Edge computing mode enables video data to be processed at the edge node, thereby reducing the bandwidth overhead. Existing edge computing-based models optimize the performance, but they still have defects in three perspectives: 1) enabling end users to control video content in a real-time format; 2) efficiently locating and transferring the user region of interest (ROI) video data in the video stream; and 3) adapting to various network conditions. To tackle these challenges, we proposed an intelligent elastic edge framework for live video analytics, known as ElasticEdge. ElasticEdge enables the interaction between the end user and the edge node. Elasticity is reflected in two perspectives: 1) the dynamic changes of user requirements and 2) the dynamic changes in network conditions. In addition, ElasticEdge transmits the video stream to the end users based on the tradeoff between the amount of video data and users’ ROI to meet various network conditions. To validate ElasticEdge, we conducted experiments to study its performance in comparison to RTFace. The experimental results show that ElasticEdge has a significant edge over RTFace in terms of data transmission. Using 1/16 reserved images, ElasticEdge saves 75% bandwidth and reduces latency by approximately 10% compared with RTFace. We also find that ElasticEdge adapts to various network conditions when streaming videos, i.e., it can reliably obtain essential information with low latency even when the network condition is poor. Hui Sun 0002, Kewei Sha |
IEEE Internet Things J. | 1 |
| 2022 | EdgeEye: A Data-Driven Approach for Optimal Deployment of Edge Video AnalyticsabstractDeep neural network (DNN)-based video processing methods are applied in mobile video analytics because of high accuracy. Edge computing is an efficient paradigm that improves the performance of mobile video analytics. However, due to the limited computing and storage resources at edge devices, deploying DNN-based video analytics at edge devices may have difficulty to meet user’s requirements in terms of accuracy, delay, power consumption, and device costs. Choosing optimal system configuration, including resources on edge devices and parameters in video stream and DNN models, can better satisfy user’s performance requirements; however, there lacks practical approaches to find such optimal configurations. In this article, we take an initial step to investigate the optimal system configuration problem, and propose a data-driven approach, EdgeEye, which first models the above problem as a combinatorial optimization problem, and then designs an algorithm to find the solutions for the optimal configuration. These models and algorithms are applied in a real-world face detection and recognition application based on two edge computing models, including edge only and edge server. Comprehensive evaluation results demonstrate that EdgeEye can find both feasible and optimal system configurations including optimal edge computing model to satisfy varying user requirements under different network conditions. Hui Sun 0002, Kewei Sha, Hong Zhong 0001 |
IEEE Internet Things J. | 1 |
| 2022 | FlexEdge: Dynamic Task Scheduling for a UAV-Based On-Demand Mobile Edge ServerabstractWith the large number of cameras deployed in smart industrial parks and smart campuses, edge devices and location-fixed edge servers are deployed near to these cameras and help transmit video streams to data center for video analytics; however, location-fixed edge servers are difficult to adapt to computation-intensive and delay-sensitive video analytics tasks in hot spot, and it is also challenging to execute tasks in natural disasters in which the infrastructure is damaged. Moreover, task migration methods are used to balance the load of edge servers caused by irregular movement of detected objects, but it results in extra data transmission overhead. Therefore, unmanned aerial vehicles (UAVs) with computing and communication resources are widely used to optimize mobile edge video analysis; however, existing solutions formulate the UAV-based lowest latency and energy consumption by jointly optimizing the task allocation strategy and UAV location to be a multiobjective optimization problem, based on which the Pareto optimum solution set, including task allocation strategies and UAV locations, can find multiple solutions but not a unique solution. It makes the solution difficult to be applied in video analytics with the UAV hover location decision-making scheme and task allocation strategy. In this article, we propose a flexible cloud-edge collaborative scheduling strategy based on a UAV namedFlexEdge. We first normalize values of execution time and energy consumption, and then convert the multiobjective optimization problem into a single-objective optimization problem by using the weighted sum of the two metrics as the optimization objective. We also proved the task allocation strategy based on execution time, energy consumption, and the UAV hover location decision-making scheme as an NP-hard problem. We propose a flexible and lightweight genetic algorithm (FGA) based on a polysomy-strengthening elitist genetic algorithm in FlexEdge to address the NP-hard problem. FlexEdge not only achieves optimal task allocation and UAV location to minimize the weighted sum of execution time and energy consumption but also provides computing resources and reliable network connection to reduce task offloading overload, which is validated by comprehensive performance evaluation. Hui Sun 0002, Bo Zhang 0111, Xiuye Zhang, Kewei Sha, Weisong Shi |
IEEE Internet Things J. | 1 |
| 2022 | HIPA: A hybrid load balancing method in SSDs for improved parallelism performance
Hui Sun 0002, Chaowei Zhang 0001, Yinliang Yue, Xiao Qin 0001 |
J. Syst. Archit. | 1 |
| 2022 | A storage computing architecture with multiple NDP devices for accelerating compaction performance in LSM-tree based KV stores
Hui Sun 0002, Yinliang Yue, Song Fu |
J. Syst. Archit. | 1 |
| 2021 | Artificial Evolution Network: A Computational Perspective on the Expansibility of the Nervous SystemabstractNeurobiologists recently found the brain can use sudden emerged channels to process information. Based on this finding, we put forward a question whether we can build a computation model that is able to integrate a sudden emerged new type of perceptual channel into itself in an online way. If such a computation model can be established, it will introduce a channel-free property to the computation model and meanwhile deepen our understanding about the extendibility of the brain. In this article, a biologically inspired neural network named artificial evolution (AE) network is proposed to handle the problem. When a new perceptual channel emerges, the neurons in the network can grow new connections to connect the emerged channel according to the Hebb rule. In this article, we design a sensory channel expansion experiment to test the AE network. The experimental results demonstrate that the AE network can handle the sudden emerged perceptual channels effectively. Youlu Xing, Hui Sun 0002, Guihuan Feng, Furao Shen, Jian Zhao 0013 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Co-Active: A Workload-Aware Collaborative Cache Management Scheme for NVMe SSDsabstractWhen it comes to NAND Flash-based solid-state disks (SSDs), cache can narrow the performance gap between user-level I/Os and flash memory. Cache management schemes impose relentless impacts on the endurance and performance of flash memory. A vast majority of existing cache management techniques adopt a passive data-update style (e.g., GCaR, LCR), thereby undermining response times in burst I/O requests-based applications11.Burst I/O requests must be served in a real-time manner. This type of I/O access pattern is prevalent in data-intensive workloads.. To address this issue, we propose a collaborative active write-back cache management scheme, called Co-Active, customized for I/O access patterns and the usage status of a flash chip. We design a hot/cold separation module to determine whether data is cold or hot in workload. When a flash chip is idle, cold and dirty data in the cache is flushed into the idle flash chip to produce clean data. To curtail cache replacement cost, clean data are preferentially evicted amid the procedure of cache replacement. A maximum write-back threshold is configured according to the level of burst I/O requests in workload. This threshold is intended to avert redundant write I/Os flushing into flash memory, thereby boosting the endurance of flash memory. The experiments are conducted to validate the advantages of Co-Active in terms of average response time, write amplification, and erase count. The findings unveil that compared with the six popular cache management schemes (LRU, CFLRU, GCaR_CFLRU, LCR, and MQSim), Co-Active (1) slashes the average response time by up to 83.89 percent with an average of 32.7 percent; (2) drives up the performance cliff degree by up to 76.4 percent with an average of 42.3 percent; and (3) improves write amplification rate by up to 60.5 percent with an average of 5.4 percent. Hui Sun 0002, Shangshang Dai, Jianzhong Huang 0001, Xiao Qin 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | LCache: Machine Learning-Enabled Cache Management in Near-Data Processing-Based Solid-State Disks
Hui Sun 0002, Shangshang Dai, Qiao Cui, Jianzhong Huang 0001 |
NPC | 1 |
| 2020 | VU: Edge Computing-Enabled Video Usefulness Detection and its Application in Large-Scale Video Surveillance SystemsabstractIn the era of smart and connected communities, video surveillance systems, which typically involve tens to thousands of cameras, have increasingly become prominent components for public safety. In current practice, when a failure occurs in a video surveillance system, the operation and maintenance teams usually spend a substantial amount of time locating and identifying the failure; hence, the fast online response cannot be guaranteed in a large-scale video surveillance system. Meanwhile, the video data that contains potential failures consumes bandwidth that could be used for useful video data. The useless video will waste the scarce bandwidth in the network and storage usage in the cloud. The emergence of edge computing is highly promising in video preprocessing with an edge camera. A video surveillance system is a killer application for edge computing. In this article, we propose an edge computing-enabled video usefulness (i.e., VU) model for large-scale video surveillance systems. We also explore its application, e.g., early failure detection and bandwidth improvement. According to the usefulness of the video data, the VU model can locate a failure and send it to end-users on the fly. In this article, our goals are threefold: 1) proposing a comprehensive VU model. To the best of our knowledge, this is the first work to explore the feasibility of the VU model and to determine VU values in a real application; 2) reducing the mean time to detection (i.e., MTTD) efficiently via edge computing-enabled fast online failure detection approaches; and 3) relieving the network bandwidth for large-scale video surveillance systems. Our experimental results demonstrate the approaches in VU model accurately detect failures that were collected from a video surveillance system with approximately 4000 cameras. The MTTD is substantially shortened by the fast online detection approaches. The video data with the worst VU values is mostly discarded to lessen overload of the network. Hui Sun 0002, Weisong Shi |
IEEE Internet Things J. | 1 |
| 2019 | Near-Data Processing-Enabled and Time-Aware Compaction Optimization for LSM-tree-based Key-Value StoresabstractWith the growing volume of storage systems, the traditional relational databases cannot reach the high performance required by big-data applications. As high-throughput alternatives to relational databases, LSM-tree-based key-value stores (KV stores in short) are confronted with degraded write performance during compaction under update-intensive workloads. To address this issue, we design and implement a time-aware compaction optimization framework for KV stores called TStore. TStore explores the near-data processing (i.e., NDP) model. It dynamically partitions compaction tasks into both host and NDP-enabled device to minimize the total time of compaction. The partitioned compaction tasks are conducted by the host and the device in parallel. The NDP-based devices exhibit low-latency, high-performance and high-bandwidth capability, thus facilitating key-value stores. TStore can not only accomplish compaction for KV stores, but also improve overall performance by removing bottleneck in compaction. Results show that the TStore with an NDP framework can achieve 3.8x and 1.9x performance improvement over LevelDB and Co-KV under the db_bench workload. In addition, the TStore-enabled KV store outperforms LevelDB and Co-KV by a factor of 3.6x and 1.9x in throughput and 72.0% and 48.9% in latency, respectively, under realistic workloads generated by YCSB. Hui Sun 0002, Jianzhong Huang 0001, Song Fu, Zhi Qiao 0001, Weisong Shi |
ICPP | 1 |
| 2019 | CalmWPC: A buffer management to calm down write performance cliff for NAND flash-based storage systems
Hui Sun 0002, Jianzhong Huang 0001, Xiao Qin 0001, Weisong Shi |
Future Gener. Comput. Syst. | 1 |
| 2019 | Collaborative Compaction Optimization System using Near-Data Processing for LSM-tree-based Key-Value Stores
Hui Sun 0002, Jianzhong Huang 0001, Weisong Shi |
J. Parallel Distributed Comput. | 1 |
| 2019 | Edge Video Analytics for Public Safety: A ReviewabstractWith the installation of enormous public safety and transportation infrastructure cameras, video analytics has come to play an essential part in public safety. Typically, video analytics is to collectively leverage the advanced computer vision (CV) and artificial intelligence (AI) to solve the four-W problem. That is to identify Who has done something (What) at a specific place (Where) at some time (When). According to the difference of latency requirements, video analytics can be applied to postevent retrospective analysis, such as archive management, search, forensic investigation and real-time live video stream analysis, such as situation awareness, alerting, and interested object (criminal suspect/missing vehicle) detection. The latter is characterized as having higher requirements on hardware resources as the sophisticated image processing algorithms under the hood. However, analyzing large-scale live video streams on the Cloud is impractical as the edge solution that conducts the video analytics on (or close to) the camera provides a silvering light. Analyzing live video streams on the edge is not trivial due to the constrained hardware resources on edge. The AI-dominated video analytics requires higher bandwidth, consumes considerable CPU/GPU resources for processing, and demands larger memory for caching. In this paper, we review the applications, algorithms, and solutions that have been proposed recently to facilitate edge video analytics for public safety. Qingyang Zhang 0001, Hui Sun 0002, Xiaopei Wu, Hong Zhong 0001 |
Proc. IEEE | 2 |
| 2019 | DLSpace: Optimizing SSD Lifetime via An Efficient Distributed Log Space Allocation StrategyabstractDue to limited numbers of program/erase cycles (i.e., P/Es) of NAND Flash, excessive out-of-place update and erase-before-write operations wear out these P/Es during garbage collections, which adversely shorten solid state disk (i.e., SSD) lifetime. The log space in NAND Flash space of an SSD performs as an updated page ′s buffer, which lowers garbage-collection frequency while reducing consumption of P/Es to extend SSD lifetime. In this article, we propose DLSpace, a novel distributed log space allocation strategy named d istributed l og space , which divides log space into block-level log space and page-level log space to significantly optimize SSD lifetime. DLSpace′s log page space is dedicated to data pages in a data block. Such log page space only buffers page-update operations in this data block; thereby the use of log blocks for postponing garbage collection delays. DLSpace is conducive to fully utilizing pages in data and log blocks to avoid erasures of blocks with free pages. Consequently, DLSpace decreases write amplification by reducing excessive valid page-rewrite and block-erase operations under random-write-intensive workloads. We carried out quantitative research on the extension of SSD lifetime by virtue of three metrics (i.e., write amplification, the number of block-erase operations, and the delay time before the first garbage collection occurring). Experimental results reveal that compared with the existing t raditional allocation strategy for l og space (i.e., TLSpace), DLSpace reduces write amplification and the number of erase operations by up to 55.2% and 64.1% to the most extent, respectively. DLSpace also extends TLSpace′s delay time of garbage collections by 73.3% to optimize SSD lifetime. Hui Sun 0002, Jianzhong Huang 0001, Xiao Qin 0001, Changsheng Xie 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | RB-Explorer: An Accurate and Practical Approach to Write Amplification Measurement for SSDsabstractA large write amplification ratio degrades the program/erase cycles (P/Es) of NAND Flashes and reduces the endurance and performance of solid state disks (SSDs). The lack of a practical way to measure write amplification for SSDs motivates us to propose a novel measuring method called RB-Explorer at the SSD level rather than the NAND Flash level. The goal of RB-Explorer is two-fold: (1) to accurately measure the write amplification of SSDs to quantify SSD endurance and (2) to study the impacts of I/O techniques on write amplification of SSDs. RB-Explorer incorporates a Ready/Busy (R/B) signal of one of the NAND Flashes in an SSD in a proposed write amplification model for SSDs with four full-parallelism levels (i.e., the channel, chip, die, and plane levels). RB-Explorer takes two steps toward measuring write amplification. First, RB-Explorer quantifies the number of page programs using the low R/B signal level, the duration of which varies with the different operation (i.e., read, program, and erase) in NAND Flash. Second, RB-Explorer measures data volume written to NAND Flashes by considering parallelisms at four levels. Data volume written to a die in a NAND Flash is obtained as a product of the number${\rm N_{p}}$of programs and page size${\rm P_{a}}$. Given the number${\rm N_{channel}}$of channels, the number${\rm N_{chip}}$of chips per channel, and the number${\rm N_{die}}$of dies per chip, one can obtain the data volume written to NAND Flashes as a product of${\rm N_{p}}, {\rm P_{a}}, {\rm N_{die}}, {\rm N_{chip}}$, and${\rm N_{channel}}$. RB-Explorer is applied to analyzing write amplification ratios of SSDs to track SSD endurance. Furthermore, we implement a real-world SSD (i.e., SSD-v) and employ a fine-tuned SSD simulator (i.e., SSDsim) to validate the accuracy of RB-Explorer. Our experimental results show that RB-Explorer improves on the accuracy of SSDsim—the state-of-the-art SSD simulator—in most tested cases. We conduct a series of measurements using micro-benchmarks and I/O traces to demonstrate how RB-Explorer may be applied to investigate SSDs. Hui Sun 0002, Xiao Qin 0001, Hong Jiang 0001, Jianzhong Huang 0001, Changsheng Xie 0001 |
IEEE Trans. Computers | 1 |
| 2014 | Exploring optimal combination of a file system and an I/O scheduler for underlying solid state disksabstractPerformance and energy consumption of a solid state disk (SSD) highly depend on file systems and I/O schedulers in operating systems. To find an optimal combination of a file system and an I/O scheduler for SSDs, we use a metric called the aggregative indicator (AI), which is the ratio of SSD performance value (e.g., data transfer rate in MB/s or throughput in IOPS) to that of energy consumption for an SSD. This metric aims to evaluate SSD performance per energy consumption and to study the SSD which delivers high performance at low energy consumption in a combination of a file system and an I/O scheduler. We also propose a metric called Cemp to study the changes of energy consumption and mean performance for an Intel SSD (SSD-I) when it provides the largest AI, lowest power, and highest performance, respectively. Using Cemp, we attempt to find the combination of a file system and an I/O scheduler to make SSD-I deliver a smooth change in energy consumption. We employ Filebench as a workload generator to simulate a wide range of workloads (i.e., varmail, fileserver, and webserver), and explore optimal combinations of file systems and I/O schedulers (i.e., optimal values of AI) for tested SSDs under different workloads. Experimental results reveal that the proposed aggregative indicator is comprehensive for exploring the optimal combination of a file system and an I/O scheduler for SSDs, compared with an individual metric. Hui Sun 0002, Xiao Qin 0001, Changsheng Xie 0001 |
J. Zhejiang Univ. Sci. C | 1 |
| 2013 | Measuring and Analyzing Write Amplification Characteristics of Solid State DisksabstractWrite amplification brings endurance challenges to NAND Flash-based solid state disks (SSDs) such as impacts upon their write endurance and lifetime. A large write amplification degrades program/erase cycles (P/Es) of NAND Flashes and reduces the endurance and performance of SSDs. The write amplification problem is mainly triggered by garbage collections, wear-leveling, metadata updates, and mapping table updates. Write amplification is defined as the ratio of data volume written by an SSD controller to data volume written by a host. In this paper, we propose a four-level model of write amplification for SSDs. The four levels considered in our model include the channel level, chip level, die level, and plane level. In light of this model, we design a method of analyzing write amplification of SSDs to trace SSD endurance and performance by incorporating the Ready/Busy (R/B) signal of NAND Flash. Our practical approach aims to measure the value of write amplification for an entire SSD rather than NAND Flashes. To validate our measurement technique and model, we implement a verified SSD (vSSD) system and perform a cross-comparison on a set of SSDs, which are stressed by micro-benchmarks and I/O traces. A new method for SSDs is adopted in our measurements to study the R/B signals of NAND Flashes in an SSD. Experimental results show that our model is accurate and the measurement technique is generally applicable to any SSDs. Hui Sun 0002, Xiao Qin 0001, Fei Wu 0005, Changsheng Xie 0001 |
MASCOTS | 1 |
| 2011 | Analysis of the File System and Block IO Scheduler for SSD in Performance and Energy ConsumptionabstractSSD (Solid State Disk) is reconsidered as the next storage device, an alternative to the HDD (Hard Disk Driver). The read/write performance and energy consumption are main aspects to the users. In our experiment, we recognize that the performance and energy consumption of SSD, based on NAND Flash, are mostly related with the file system and block I/O scheduler. In order to gain higher performance and lower energy consumption, we test the different combination of file system and scheduler under workload simulator, File bench, using three kinds of commercial SSDs. According to the different combination of file system and block I/O, we analyze the performance parameter, IOPS, and energy consumption parameter, POWER, under some special workload. Lastly, we present a parameter, aggregative indicator (AI), to evaluate the overall characteristic of some combination of file system and block I/O, which synthesizes IOPS and POWER. It is to find a better combination for special workload. In the experiment, the combination of extent file system (ext2 or ext3) and CFQ expresses better more value of the aggregative indicator than others. Hui Sun 0002, Fei Wu 0005, Changsheng Xie 0001 |
APSCC | 1 |