EDBT 2026 Demo / reviewers in the wild / expert
Litong You
dblp:234/1602
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shiro: Efficient and Accurate In-Storage Data Lifetime Separation for nand Flash SSDsabstractThe log-structured nature of NAND flash storage necessitates garbage collection in SSDs. Garbage collection (GC) is a major source of runtime write amplification (WA), leading to faster device wear out and interference with host I/Os. The key to mitigating this problem is separating data by lifetime so that data in the same flash block are invalidated within temporal proximity. For higher lifetime prediction accuracy and adaptibility, prior works proposed using machine learning algorithms for data separation. However, existing learning-based solutions perform data lifetime prediction at the host side, leading to several drawbacks. First, host-side prediction does not have knowledge of the internal data movement inside the SSD during GC, and thus fails to leverage the opportunity to further separate GC writes, resulting in suboptimal WA reduction in the long term. Second, performing prediction at the host significantly prolongs the I/O critical path and consumes host resources that could otherwise be used for serving user applications. We present Shiro, a holistic FTL design that performs instorage data separation for both user writes and GC writes for maximal long-term WA reduction. For user writes, Shiro uses a sequence model to accurately predict data lifetime by learning lifetime distribution from long historical access patterns. For GC writes, Shiro incorporates a reinforcement learning-assisted page migration strategy that takes direct feedback from longterm WA to further improve data separation efficacy. To address the challenges posed by performing fine-grained and real-time machine learning decisions inside the resource-constrained SSD, we propose a suite of enabling techniques to keep computation and storage overhead low. Extensive evaluation of Shiro on real-world traces shows that Shiro can deliver 29 WA compared with conventional FTL and state-of-the-art instorage data separation schemes. Furthermore, thanks to lower data migration overhead during GC, Shiro achieves significantly higher steady-state I/O performance. Penghao Sun, Shengan Zheng, Litong You, Wanru Zhang, Ruoyan Ma, Feng Zhu 0024, Linpeng Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Learning-based Data Separation for Write Amplification Reduction in Solid State DrivesabstractGarbage collection in SSDs causes write amplification. The key to mitigating this problem is separating data by lifetime. Prior works proposed using machine learning to accurately predict data lifetime but prediction is performed at the host side, burdening the host storage stack. We present PHFTL, a practical, holistic FTL design with device-side learning-based data separation. The machine learning model in PHFTL accurately and adaptively predicts the lifetime of every written page. A suite of enabling techniques are introduced to keep computation and storage overhead low. Extensive evaluation of PHFTL demonstrates superiority over state-of-the-art and feasibility on real hardware. Penghao Sun, Litong You, Shengan Zheng, Wanru Zhang, Ruoyan Ma, Guanzhong Wang, Feng Zhu 0024, Linpeng Huang |
DAC | 2 |
| 2021 | JPDHeap: A JVM Heap Design for PM-DRAM MemoriesabstractReal-world e-commerce systems need large cache capacities. Persistent memory (PM) can be employed to enlarge JVMs’ cache capacities, meanwhile they incur heavy write slowdowns and garbage collection overheads. This paper proposes JPDheap, a JVM heap design for PM-DRAM memories. A JPDheap is composed of a standard Java heap on DRAM and another heap on PM. The core insight is to separate heap objects and store them on DRAM or PM, allowing objects to be accessed much more efficiently. Our evaluation shows that JPDheap outperforms state-of-the-art heap designs by up to 115.96% in increasing applications’ throughput and by up to 87.03% in decreasing the average latency. Litong You, Tianxiao Gu, Shengan Zheng, Jianmei Guo, Sanhong Li, Yuting Chen 0001, Linpeng Huang |
DAC | 1 |
| 2019 | JDap: Supporting in-memory data persistence in javascript using Intel's PMDK
Litong You, Qipeng Zhang, Tianyou Li, Chen Li 0009, Yuting Chen 0001, Linpeng Huang |
J. Syst. Archit. | 1 |
| 2018 | Forca: Fast and Atomic Remote Direct Access to Persistent MemoryabstractFor promising performance boost, recent trends of modern data centers tend to use Persistent Memory (PM) as storage and utilize the one-sided feature of Remote Direct Memory Access (RDMA) to directly access PM for I/O requests. However, accessing PM through one-sided RDMA faces two challenges: one-sided data races and remote data crash consistency. Existing systems that employ server-bypass mechanism either abandon server-bypass write to avoid the challenges at the cost of extra server loads, or support both server-bypass read/write through inefficient concurrency control without ensuring remote crash consistency. In this paper, we propose a novel server-bypass RDMA-to-PM framework, named Forca, to provide high concurrency and guarantee remote crash consistency at the same time. In Forca, we design an optimized log-structured mechanism to eliminate in-place updates, removing race conditions of one-sided RDMA and providing atomicity for each update simultaneously. We implement Forca as a generic module to support server-bypass RDMA to PM, and conduct experiments on a Forca-based key-value store called ForcaKV. The experiments show that Forca achieves higher throughput than state-of-the-art techniques by up to 1.4x under high concurrency scenarios, and its crash consistency mechanism incurs only 14.7%-19.1% time overhead. Haixin Huang, Kaixin Huang, Litong You, Linpeng Huang |
ICCD | 3 |