Mengting Lu

dblp:245/6113 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0003-2333-2123ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2024 TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
abstract
Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform with the operating system and CXL/RDMA-based remote memory pools in two key areas. Firstly, TrEnv introduces repurposable sandboxes, which can be shared across different functions and hence, substantially decrease the overhead associated with creating isolation sandboxes. Secondly, it augments the OS with "memory templates" that enable rapid restoration of function states stored on remote memory. These innovations allow TrEnv to facilitate rapid transitions between instances of different functions and enable memory sharing across multiple nodes. Our evaluations using a variety of representative and real-world workloads demonstrate that TrEnv can initiate a container within 10 milliseconds, achieving up to a 7× speedup in P99 end-to-end latency and reducing memory usage by 48% on average compared to state-of-the-art on-demand restoring systems.
Teng Ma 0006, Zheng Liu 0022, Sixing Lin, Kang Chen 0001, Jinlei Jiang, Xia Liao, Yingdi Shan, Mengting Lu, Tao Ma 0006, Haifeng Gong, Yongwei Wu 0001
SOSP11
2023 CostFM: A High Cost-Performance Fingerprint Management Mechanism for Shared SSDs
abstract
As the storage and I/O demand of contemporary computer systems grow, SSDs with deduplication are widely deployed as shared storage devices to provide high performance in cloud platforms where diverse tenant workloads are collocated. However, existing global and fixed fingerprint management schemes are inefficient in the multi-tenant environment. Contention for in-memory fingerprint cache, embedded in DRAM to speed up fingerprint lookups, results in decreased cache utilization, and abundant fingerprint lookups into the backend flash memory degrade the system performance. Besides, a fixed cache replacement policy for all the tenants fails to capture the diverse characteristics, further increasing the lookup overhead.This paper introduces CostFM, a novel fingerprint management mechanism highlighted by two notable features. First, it provides a benefit-aware cache allocation scheme, which estimates the cache hit gains of each tenant and uses dynamic programming to find the optimal cache allocation scheme, avoiding the cache contention for the fingerprint resource and improving cache utilization. Second, it applies a user-based policy model that captures access characteristics to select the suitable cache management policy for each tenant, boosting overall cache efficiency. Extensive experimental results show that CostFM decreases the average request latency by up to 30%, and it also reduces the fingerprint writes into the flash memory by up to 4.7x.
Mengting Lu, Fang Wang 0001, Wenpeng He
ICCD2
2023 PH-ORAM: An efficient persistent ORAM design for hybrid memory systems
Wenpeng He, Dan Feng 0001, Fang Wang 0001, Baoquan Li, Mengting Lu
J. Syst. Archit.5
2022 IRO: Integrity-Reliability enhanced Ring ORAM
Wenpeng He, Dan Feng 0001, Fang Wang 0001, Mengting Lu
J. Syst. Archit.5
2022 SPOPB: Reducing solid state drive write traffic for flash-based key-value caching
abstract
Abstract Flash‐based key‐value (KV) caching has received increasing attention in recent years with the advantages of flash‐based solid state drives (SSDs) in capacity and cost. By caching most data in SSD, the caching system can eliminate lots of time‐consuming requests to back‐end data stores to provide low‐latency services. To adapt to the unique technical constraints of flash memory, flash‐based KV caching adopts a slab‐based log‐structured management scheme in which the slab is the basic storage unit, and uses a small memory space as a write buffer to eliminate small random writes to SSD for consistent performance and increased lifetime of SSD. However, we have observed that under update‐intensive workloads with strong temporal locality, the slab‐based management in flash‐based KV caching introduces substantial SSD write traffic because of indistinguishable SSD flushing of hot items in slabs, which shortens the SSD lifetime and degrades the performance with increased erase operations. In this article, we first analyze the SSD write traffic in the flash‐based KV caching, and then propose a novel slab popularity‐based storage management scheme‐SPOPB, to extend SSD lifetime and improve performance. Our scheme identifies hot items using a self‐adaptive threshold to reorganize and classify slabs with both the hotness and size of items. Then SPOPB filters and retains the popular slabs containing hot items in the write buffer with redesigned replacement policy to reduce the SSD write traffic. Our experiments show that our design can effectively reduce the SSD write traffic by 63.6%, the erase counts by 55.6%, and improve the performance by 42%.
Zongwei Li 0001, Dan Feng 0001, Yuchong Hu, Mengting Lu
Softw. Pract. Exp.4
2022 EDC: An Elastic Data Cache to Optimizing the I/O Performance in Deduplicated SSDs
abstract
Data deduplication is widely deployed in solid-state drives (SSDs) to improve the storage space utilization and alleviate the endurance issue. However, deduplication increases the degree of fragmentation and the access contention, significantly hurting the system performance. First, the fragmentation at the storage medium level in SSD is closely related to the internal parallelism, and the degree of fragmentation increases as the degree of the parallelism decreases. Deduplication removes the duplicate parts of the write sequences, decreasing the read parallelism. Thus, the degree of fragmentation is increased, eventually degrading the read performance. Second, the uneven distribution of the highly referenced data increases the access contention. This increased access contention prolongs the queueing time, further degrading the system performance. Motivated by our observations, we propose an elastic data cache (EDC) to improve the I/O performance in the deduplicated SSD. EDC redesigns the built-in DRAM-based data cache, tracks the popular and highly-referenced data, and maintains them in the cache. To fulfill the novel data cache, EDC changes the request process. The read requests that access the fragments in the flash memory are mostly performed in the fast-speed data cache, alleviating the negative impact of fragmentation on the read performance. The reduced read accesses to the flash memory also ease the access contention, which improves the system performance significantly. Extensive experimental results validate the efficiency of EDC, showing that it effectively improves the read performance and the write performance by up to 79% and 85% on average, respectively, in the deduplicated SSD.
Mengting Lu, Fang Wang 0001, Wenpeng He
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 A Read-leveling Data Distribution Scheme for Promoting Read Performance in SSDs with Deduplication
abstract
Deduplication, as a space-saving technology, is widely deployed in the flash-based storage systems to address the capacity and endurance limitations of flash devices. In this paper, we find that deduplication changes the physical data layout, which raises the chances of the uneven read distribution. This uneven read distribution not only increases the access contention but also deteriorates the read parallelism, thus leading to the read performance degradation. To solve this issue, we propose an efficient read-leveling data distribution scheme (RLDDS), which scatters the highly-duplicated data into different parallel units, to improve the read performance for SSDs with deduplication for access-intensive workloads. RLDDS writes data into a parallel unit with lower potential read-hotness to balance the read distribution among all the parallel units. Extensive experimental results show that RLDDS effectively improves the read performance by up to 21.61% compared to deduplication with the conventional dynamic data allocation scheme. Additional benefits of RLDDS include the promoted write performance (up to 23.69%) in access-intensive workloads and the overall system performance improvement (up to 18.22%) with the same write traffic reduction.
Mengting Lu, Fang Wang 0001, Dan Feng 0001, Yuchong Hu
ICPP1