VLDB 2026 Research / reviewers in the wild / expert
Dingding Li
dblp:40/99
· DBLP profile ↗
21ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-9092-9814ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 10 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAIKA: Machine Learning-Assisted In-Kernel APU AccelerationabstractThe integration of machine learning (ML) into OS kernels is severely hampered by the high latency of offloading to discrete GPUs (dGPUs), where data transfers across the PCIe bus can consume over 93% of the total execution time. This paper argues that for many latency-sensitive kernel tasks, the solution is not a more powerful dGPU but a fundamental shift to an I/O-efficient architecture: the integrated GPU (iGPU) found in modern APUs. Haoming Zhuo, Dingding Li, Ronghua Lin, Yong Tang 0001 |
ASPLOS (2) | 2 |
| 2024 | A Sorted and Dynamic Graph Storage System of the Hybrid Memory ArchitectureabstractGraph data is becoming dynamic and large-scale, demanding high-performance and large-capacity graph storage. Therefore, due to the performance approaching DRAM and the larger capacity than DRAM, persistent memory (PM) has been adopted in large-scale dynamic graph storage systems. However, existing PM-based dynamic graph storage systems have issues, especially PM write amplification caused by the unsorted data structure used to store edges. To improve this issue, we propose a PM-based dynamic graph storage system, HDGraph, using sorted data structure to store edges on DRAM-PM hybrid memory architecture. To better adapt the sorted data structure on PM, HDGraph employs edge buffering on DRAM, merging small writes to reduce write amplification in PM. Moreover, HDGraph also triggers buffer flushing based on a heat evaluating strategy to alleviate DRAM space pressure. Finally, HDGraph maintains a buffering log in PM for edge-level data consistency, enabling quick recovery after a crash. Experimental results show that HDGraph achieves 1.09× to 1.52× higher edge ingestion performance, compared with the only PM-based dynamic graph storage system XPGraph, which use unsorted data structure to store edges. Yang Ao, Yutong Deng, Dingding Li, Yong Tang 0001 |
ISPA | 3 |
| 2024 | A LLM Checkpointing System of the Hybrid Memory ArchitectureabstractThe training process of large language models (LLMs) sometimes may raise an error, often necessitating the usage of checkpointing mechanisms to save training states periodically. However, this can lead to a waste of training time and impact performance due to frequent I/O operations. Intel Optane Data Center Persistent Memory Module (DCPMM) is a novel memory device that combines non-volatility, high bandwidth, and low latency. Intuitively, deploying LLMs on DCPMM could enhance I/O performance, but the DCPMM cannot fully adapt to the current deep learning frameworks and the read/write granularity mismatch of DCPMM needs to be resolved. To this end, we propose PMCKPT, a checkpointing scheme customized for hybrid memory architectures (DRAM-PM). PMCKPT provides a storage engine to mitigate the impact of frequent I/O operations on training. Furthermore, PMCKPT ensures data persistence and consistency through a persistence scheme based on the persistent memory development kit (PMDK). Finally, it also incorporates a checkpointing management strategy to handle historical version checkpointing and optimize storage management. We implement PMCKPT in ChatGLM and the experimental results demonstrate that PMCKPT can effectively reduce training time by about 34% compared to traditional schemes which use SSDs. Gengbin Chen, Dingding Li, Yong Tang 0001 |
ISPA | 3 |
| 2024 | DMA-Assisted I/O for Persistent MemoryabstractModern local persistent memory (PM) file systems often rely on CPU-based memory copying for data transfer between DRAM and PM, resulting in significant CPU resource consumption. While some nascent systems explore DMA (direct memory access) as an alternative for improved efficiency, the intricacies and trade-offs remain obscure. This paper investigates the feasibility of DMA for PM I/O and argues that it is not a straightforward replacement for CPU-based methods. Two key limitations hinder the direct adoption: poor performance for small data and limited bandwidth. To relieve these issues, we propose PM-DMA, a novel I/O mechanism that leverages the strengths of both CPU and DMA. It incorporates three key components: (1) L-Switch, seamlessly switches between CPU and DMA modes based on workload characteristics, maximizing performance; (2) D-Pool, reduces DMA setup overhead, improving responsiveness; (3) P-Mode, allows servicing requests through multiple channels, even hybrid CPU-DMA ones, for enhanced throughput. We implemented PM-DMA on two well-known PM file systems, NOVA and WineFS, utilizing Intel I/OAT technology. Our experimental results demonstrate substantial CPU consumption reductions across diverse workloads. Notably, under heavy load, PM-DMA delivers up to a$10.4\times$performance improvement. Dingding Li, Mianxiong Dong, Kaoru Ota |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | PM-Migration: A Page Placement Mechanism for Real-Time Systems with Hybrid Memory Architecture
Lidang Xu, Gengbin Chen, Dingding Li, Haoyu Luo |
ICA3PP (5) | 3 |
| 2022 | A DMA-based Swap Mechanism of Hybrid Memory SystemabstractTypical applications of smart cities, such as smart public services, require a large memory footprint to store user data and facilitate the responsive results of user queries, thus inevitably activating the memory swap mechanism between memory and storage to expand the capacity of main memory. Frequent page swapping can cause performance interference for hard real-time operating systems such as SylixOS. In a hybrid memory architecture, namely the novel persistent memory (PM) alongside the conventional DRAM, the swap mechanism often uses the PM to act as the swap partition and executes memory copying to transfer the data between DRAM and PM, resulting in frequent I/O operations and high CPU consumption. Eventually, the memory performance is sub-optimal. By leveraging a general DMA technology of memory-to-memory (M2M), namely Intel I/OAT, we propose PM-Swap, a swap mechanism without heavy CPU consumption. PM-Swap further contains three techniques: (1) a new memory reclamation algorithm based on instruction sampling and page awareness, which reduces the unnecessary swap operations; (2) according to the data size, a switching strat-egy selects the suitable swapping path between the original CPU and the DMA, to maintain reasonable memory performance; (3) bulk transferring is employed for improving the overall throughput of the page swapping. We implement PM-Swap in a stable Linux kernel (5.17.9). The experimental results show that PM-Swap can decrease CPU overhead by more than 39% and increase page swapping bandwidth by up to$1.76\times$. Lidang Xu, Dingding Li, Haoyu Luo |
MSN | 3 |
| 2022 | FSbrain: An intelligent I/O performance tuning system
Yong Tang 0001, Ronghua Lin, Dingding Li, Yuguo Li, Deze Zeng |
J. Syst. Archit. | 3 |
| 2021 | A Lightweight Asynchronous I/O System for Non-volatile Memory
Jiebin Luo, Dingding Li, Haoyu Luo, Deze Zeng |
ICA3PP (2) | 3 |
| 2021 | IObrain: An Intelligent Lightweight I/O Recommendation System based on Decision TreeabstractThe basic I/O operations of a system can be categorized as two distinct modes: synchronous (sync) I/O and asynchronous (async) I/O, whose performance varies on the system statues, workloads and storage devices. Appropriately applying I/O modes is critical to the system performance. However, the I/O access of diverse applications in a server, especially in a cloud, is volatile and irregular. As a result, this can lack a flexible and adaptive I/O modes, leading to the sub-optimal I/O performance. To tackle this problem, in this paper, we propose IObrain, an intelligent I/O mode recommendation system, which can adopt the appropriate I/O mode in a dynamic and self-adaptive manner according to both application needs and system statuses. IObrain first trains a lightweight recommendation model with decision tree. Then, a query hook is interposed into the storage engine to intercept the read/write operations from upper application. In this way, IObrain queries the recommendation model first before executing a read/write operation to find the right I/O mode. In addition, two techniques, called inference cache and gRPC bridge, are proposed to reduce the inherent query latency. We practically implement IObrain and verify the advantage of IObrain based on the prototype system. The experimental results show that, compared to existing approach, IObrain improves the I/O performance by up to 1.33× with mild running costs. Yuguo Li, Junlang Huang, Dingding Li, Yong Tang 0001, Deze Zeng |
ICPADS | 5 |
| 2021 | RAMCI: a novel asynchronous memory copying mechanism based on I/OAT
Zhenke Chen, Dingding Li, Hai Liu 0006, Yong Tang 0001 |
CCF Trans. High Perform. Comput. | 2 |
| 2021 | Towards High-Efficient Transaction Commitment in a Virtualized and Sustainable RDBMSabstractThe relational database service in cloud usually achieves energy efficiency by using virtualization technology, in which it consolidates multiple independent database systems into a single physical machine while enforcing the hardware-level isolation among them. However, the disk I/O performance is inevitably hurt due to the resource contention on the shared device. We propose VMSQL, a novel disk I/O model for the virtualized relational database management system (RDBMS). VMSQL has two innovations over the original disk model of virtualized database systems. First, VMSQL enforces the synchronous operation in guest operating system to handle with the transaction commitment. Due to its simplicity, a portion of CPU cycles is decoupled from I/O buffer management and then used to serve the upcoming requests, thereby improving their response times. Second, in host system, VMSQL asynchronizes the storage path of transactions which are committed from the different co-located guest databases. An obvious advantage of this procedure is that systems can apply host-level improvements into the disk I/O performance of virtualized RDBMS, relieving the random I/O and enhancing the throughput of whole system. We implement a prototype of this Sync-Async model in QEMU-KVM hypervisor, in which the InnoDB engine is deployed in the guest operating system. Extensive experiments are conducted to verify its advantages and the results are positive without any loss of ACID-compliance. In the meanwhile, VMSQL incurs moderate overhead at the hypervisor layer. Dingding Li, Kaoru Ota, Mianxiong Dong, Yong Tang 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2018 | SEER-MCache: A Prefetchable Memory Object Caching System for IoT Real-Time Data ProcessingabstractMemory object caching systems, such as Memcached and Redis, have been proved to be a simple and high-efficient middleware for improving the performance of Internet of Things (IoT) devices querying the database in cloud. However, its performance guarantee is built on the fact that the target data, queried by the IoT device, will be accessed many times and hit in the caching system. Therefore, when database system is handling the unrepeated IoT queries, it usually presents the suboptimal performance, which greatly impairs the efficiency of real-time data processing on IoT devices. To improve this issue, we propose Seer-MCache, the memory object caching system with a smart prefetching (read-ahead) function, to fill up the caching system with the desired data before the intensive IoT queries arriving. Seer-MCache includes a set of rules to launch the specific behaviors of read-head. These rules are able to be customized according to the workload characteristics and system load. We implement a prototype system in Redis (caching layer) and MySQL server (database system). Extensive experiments are conducted to verify the effectiveness of Seer-MCache, the results show that Seer-MCache can improve the performance of read-intensive workload up to 61% (39.5% in average). Meanwhile, the cost of the read-ahead behavior is moderate and controllable. Dingding Li, Mianxiong Dong, Yanting Yuan, Kaoru Ota, Yong Tang 0001 |
IEEE Internet Things J. | 1 |
| 2017 | Citation Based Collaborative Summarization of Scientific Publications by a New Sentence Similarity Measure
Chengzhe Yuan, Dingding Li, Jia Zhu 0003, Yong Tang 0001, Shahbaz Hassan Wasti, Chaobo He, Hai Liu 0006, Ronghua Lin |
CollaborateCom | 2 |
| 2016 | Speeding Up Virtualized Transaction Logging with vTransabstractIn a virtualized environment, when multiple co-located relational database engines update their data files on a shared storage device simultaneously, the transaction log files, which be scattered within different guest image files, introduce the random logging I/O easily. To cope with this issue, we propose vTrans, a novel I/O driver of virtualized transaction log file. Generally, vTrans uses the split I/O driver design. In each guest operating system a front-end driver is added to identify the transaction logging semantic from specific database engine. At the hypervisor layer a dedicated back-end driver consolidates all transaction logging I/Os from guest font-end drivers and then persist them into a consecutive data area on the shared storage device, resulting in the relatively sequential logging I/O. We implement vTrans in a QEMU system which deployed MySQL InnoDB database engine. The experimental result shows that vTrans can effectively improve the performance of random logging I/O in a virtualized system. Yong Tang 0001, Lingxiao Chen, Dingding Li |
ICPADS | 4 |
| 2016 | Writeback throttling in a virtualized system with SCM
Dingding Li, Xiaofei Liao, Hai Jin 0001, Yong Tang 0001, Gansen Zhao |
Frontiers Comput. Sci. | 1 |
| 2015 | Improving write amplification in a virtualized and multimedia SSD system
Dingding Li, Hai Jin 0001, Xiaofei Liao, Jia Yu 0010 |
Multim. Tools Appl. | 1 |
| 2013 | A Performance Optimization Mechanism for SSD in Virtualized EnvironmentabstractApplications in the cloud computing era have the emergent requirement of fast I/O support. Compared with a hard drive disk, a solid state disk (SSD) has low delay, low energy consumption, high throughput and other advantages. However, the semantics of an SSD cannot be recognized by current virtual machine monitors. The trim instruction, which plays an important role in space management of an SSD, cannot be passed to the underlying SSD device in a virtualized environment. So, how to bridge the semantics gap between the application layer and virtualization layer for SSD devices should be an important problem. We propose Vtrim to solve the above problem in this paper. Vtrim monitors the operations in the virtual machine and sends the SSD semantics to the Domain 0 immediately. And then the semantics are translated into the operations of Domain 0, which can trigger the SSD's local instructions. To improve the write performance in multiple guest operating systems, we set a Vtrim cache to buffer all instructions from the guests and flush them into an SSD in a well-scheduled way. Experiment results in the para-virtualization environments with Vtrim show that the random write performance is improved by even up to 100% and the average response time is reduced by up to 40%. Xiaofei Liao, Hai Jin 0001, Jia Yu 0010, Dingding Li |
Comput. J. | 4 |
| 2013 | Improving disk I/O performance in a virtualized system
Dingding Li, Hai Jin 0001, Xiaofei Liao, Yu Zhang 0027, Bing Bing Zhou |
J. Comput. Syst. Sci. | 1 |
| 2013 | Bias Correction in a Small Sample from Big DataabstractThis paper discusses the bias problem when estimating the population size of big data such as online social networks (OSN) using uniform random sampling and simple random walk. Unlike the traditional estimation problem where the sample size is not very small relative to the data size, in big data, a small sample relative to the data size is already very large and costly to obtain. We point out that when small samples are used, there is a bias that is no longer negligible. This paper shows analytically that the relative bias can be approximated by the reciprocal of the number of collisions; thereby, a bias correction estimator is introduced. The result is further supported by both simulation studies and the real Twitter network that contains 41.7 million nodes. Jianguo Lu, Dingding Li |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | A New Disk I/O Model of Virtualized Cloud EnvironmentabstractIn a traditional virtualized cloud environment, using asynchronous I/O in the guest file system and synchronous I/O in the host file system to handle an asynchronous user disk write exhibits several drawbacks, such as performance disturbance among different guests and consistency maintenance across guest failures. To improve these issues, this paper introduces a novel disk I/O model for virtualized cloud system called HypeGear, where the guest file system uses synchronous operations to deal with the guest write request and the host file system performs asynchronous operations to write the data to the hard disk. A prototype system is implemented on the Xen hypervisor and our experimental results verify that this new model has many advantages over the conventional asynchronous-synchronous model. We also evaluate the overhead of asynchronous I/O at host, which is brought by our new model. The result demonstrates that it enforces little cost on host layer. Dingding Li, Xiaofei Liao, Hai Jin 0001, Bing Bing Zhou, Qi Zhang 0009 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2010 | Estimating deep web data source size by capture-recapture method
Jianguo Lu, Dingding Li |
Inf. Retr. | 2 |