VLDB 2026 Research / reviewers in the wild / expert
Lei Liu 0037
dblp:21/2715-37
· DBLP profile ↗
20ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-4854-7382ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?abstractMaking it intelligent is a promising way in System/OS design. This article proposes OSML+, a new ML-based resource scheduling mechanism for co-located cloud services. OSML+ intelligently schedules the cache and main memory bandwidth resources at the memory hierarchy and the computing core resources simultaneously. OSML+ uses a multi-model collaborative learning approach during its scheduling and thus can handle complicated cases, e.g., avoiding resource cliffs, sharing resources among applications, enabling different scheduling policies for applications with different priorities, and so on. OSML+ can converge faster using ML models than previous studies. Moreover, OSML+ can automatically learn on the fly and handle dynamically changing workloads accordingly. Using transfer learning technologies, we show our design can work well across various cloud servers, including the latest off-the-shelf large-scale servers. Our experimental results show that OSML+ supports higher loads and meets QoS targets with lower overheads than previous studies. Xinglei Dou, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Storage | 2 |
| 2025 | Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data TypeabstractThe quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator. Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao |
MICRO | 5 |
| 2025 | Exploiting intra-chip locality for multi-chip GPUs via two-level shared L1 cache
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Xiaojian Liao |
J. Syst. Archit. | 4 |
| 2025 | LarQucut: A New Cutting and Mapping Approach for Large-sized Quantum Circuits in Distributed Quantum Computing (DQC) EnvironmentsabstractDistributed quantum computing (DQC) is a promising way to achieve large-scale quantum computing. However, mapping large-sized quantum circuits in DQC is a challenging job; for example, it is difficult to find an ideal cutting and mapping solution when many qubits, complicated qubit operations, and diverse QPUs are involved. In this study, we propose LarQucut, a new quantum circuit cutting and mapping approach for large-sized circuits in DQC. LarQucut has several new designs. (1) LarQucut can have cutting solutions that use fewer cuts, and it does not cut a circuit into independent sub-circuits, therefore reducing the overall cutting and computing overheads. (2) LarQucut finds isomorphic sub-circuits and reuses their execution results. So, LarQucut can reduce the number of sub-circuits that need to be executed to reconstruct the large circuit's output, reducing the time spent on sampling the sub-circuits. (3) We design an adaptive quantum circuit mapping approach, which identifies qubit interaction patterns and accordingly enables the best-fit mapping policy in DQC. The experimental results show that, for large circuits with hundreds to thousands of qubits in DQC, LarQucut can provide a better cutting and mapping solution with lower overall overheads and achieves results closer to the ground truth. Xinglei Dou, Lei Liu 0037, Zhuohao Wang |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | An Intelligent Scheduling Approach on Mobile OS for Optimizing UI Smoothness and PowerabstractMobile devices need to respond quickly to diverse user inputs. The existing approaches often heuristically raise the CPU/GPU frequency according to the empirical rules when facing burst inputs and various changes. Although doing so can be effective sometimes, the existing approaches still need improvements. For instance, raising processors’ frequency can lead to high power consumption when the frequency is over-provisioned or fail to meet user demands when the frequency is under-provisioned. To this end, we propose MobiRL, a reinforcement learning-based scheduler for intelligently adjusting the CPU/GPU frequency to satisfy user demands accurately on mobile systems. MobiRL monitors the mobile system status and autonomously learns to optimize UI smoothness and power consumption by conducting CPU/GPU frequency-adjusting actions. The experimental results on the latest delivered smartphones show that MobiRL outperforms the widely used commercial scheduler on real devices—reducing the frame drop rate by 4.1% and reducing power consumption by 42.8%, respectively. Moreover, compared with a study using Q-Learning for CPU frequency scheduling, MobiRL achieves up to a 2.5% lower frame drop rate and reduces power consumption by 32.6%, respectively. Our approach has been deployed in mobile phone products. Xinglei Dou, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | QuCloud+: A Holistic Qubit Mapping Scheme for Single/Multi-programming on 2D/3D NISQ Quantum ComputersabstractQubit mapping for NISQ superconducting quantum computers is essential to fidelity and resource utilization. The existing qubit mapping schemes meet challenges, e.g., crosstalk, SWAP overheads, diverse device topologies, etc., leading to qubit resource underutilization and low fidelity in computing results. This article introduces QuCloud+, a new qubit mapping scheme that tackles these challenges. QuCloud+ has several new designs. (1) QuCloud+ supports single/multi-programming quantum computing on quantum chips with 2D/3D topology. (2) QuCloud+ partitions physical qubits for concurrent quantum programs with the crosstalk-aware community detection technique and further allocates qubits according to qubit degree, improving fidelity, and resource utilization. (3) QuCloud+ includes an X-SWAP mechanism that avoids SWAPs with high crosstalk errors and enables inter-program SWAPs to reduce the SWAP overheads. (4) QuCloud+ schedules concurrent quantum programs to be mapped and executed based on estimated fidelity for the best practice. Experimental results show that, compared with the existing typical multi-programming study [ 12 ], QuCloud+ achieves up to 9.03% higher fidelity and saves on the required SWAPs during mapping, reducing the number of CNOT gates inserted by 40.92%. Compared with a recent study [ 30 ] that enables post-mapping gate optimizations to further reduce gates, QuCloud+ reduces the post-mapping circuit depth by 21.91% while using a similar number of gates. Lei Liu 0037, Xinglei Dou |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | iSwap: A New Memory Page Swap Mechanism for Reducing Ineffective I/O Operations in Cloud EnvironmentsabstractThis article proposes iSwap , a new memory page swap mechanism that reduces the ineffective I/O swap operations and improves the QoS for applications with a high priority in cloud environments. iSwap works in the OS kernel. iSwap accurately learns the reuse patterns for memory pages and makes the swap decisions accordingly to avoid ineffective operations. In the cases where memory pressure is high, iSwap compresses pages that belong to the latency-critical (LC) applications (or high-priority applications) and keeps them in main memory, avoiding I/O operations for these LC applications to ensure QoS, and iSwap evicts low-priority applications’ pages out of main memory. iSwap has a low overhead and works well for cloud applications with large memory footprints. We evaluate iSwap on Intel x86 and ARM platforms. The experimental results show that iSwap can significantly reduce ineffective swap operations (8.0%–19.2%) and improve the QoS for LC applications (36.8%–91.3%) in cases where memory pressure is high, compared with the latest LRU-based approach widely used in modern OSes. Zhuohao Wang, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | ATA-Cache: Contention Mitigation for GPU Shared L1 Cache With Aggregated Tag ArrayabstractTo fully exploit the locality of GPU applications, the GPU shared L1 cache architecture, which shares L1 cache among multiple GPU cores, is a promising architecture while still suffering from high resource contentions. We present a GPU shared L1 cache architecture with an aggregated tag array that minimizes the L1 cache contentions and takes full advantage of inter-core locality. The key idea is to decouple and aggregate the tag arrays of multiple L1 caches so that the cache requests can be compared with all tag arrays in parallel to probe the replicated data in other caches. The GPU caches are only accessed by other GPU cores when replicated data exists, filtering out unnecessary cache accesses that cause high resource contentions. We also develop a two-level thread-block scheduling policy adapted for the shared L1 cache architecture to maximize the available locality. The experimental results show that GPU performance can be improved by 14.5% on average for applications with a high inter-core locality. Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Hao Liu 0107 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning Approach
Lei Liu 0037, Xinglei Dou, Yuetao Chen |
FAST | 1 |
| 2023 | CFIO: A conflict-free I/O mechanism to fully exploit internal parallelism for Open-Channel SSDs
Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin |
J. Syst. Archit. | 4 |
| 2023 | EBIO: An Efficient Block I/O Stack for NVMe SSDs With Mixed WorkloadsabstractWith the advent of high-performance nonvolatile memory express (NVMe) SSD, the overhead caused by the storage software stack becomes a significant bottleneck for exploiting the potential of NVMe SSD. Recent I/O isolation approaches eliminate CPU switching, I/O interference, and lock contention for I/O queues by pinning I/O threads in isolated and dedicated I/O paths. However, they degrade the overall performance of mixed workloads with heterogeneous I/O demands. The I/O-intensive workloads issue multiple I/O requests and quickly fill up their dedicated I/O queues, resulting in I/O wait. On the contrary, the non-I/O-intensive workloads cannot deliver enough I/O requests to saturate their associated I/O queues. Moreover, frequent allocations and deallocations of I/O request objects expose a significant overhead for I/O-intensive workloads. In this article, we propose EBIO, an efficient block I/O stack for NVMe SSD, to improve the overall performance of mixed workloads. EBIO contains an on-demand queue management strategy (ODQM) and a reusable I/O management strategy (RERM). Specifically, ODQM eliminates I/O wait by dynamically adding queues for I/O-intensive workloads and leverages I/O generation time to guarantee strong sequential consistency and fairness in scheduling I/O requests. RERM reduces the overhead caused by repeatedly allocating I/O request objects for I/O-intensive workloads by reusing the allocated objects. Experimental results show that, compared to the state-of-the-art approaches, EBIO improves input/output operations per second by up to 16.44% and reduces I/O latency by up to 32.76%. Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 7 |
| 2022 | Publisher Correction: Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 7 |
| 2022 | Machine-learning-based cache partition method in cloud environment
Jiefan Qiu, Zonghan Hua, Lei Liu 0037, Mingsheng Cao 0001, Dajiang Chen |
Peer-to-Peer Netw. Appl. | 3 |
| 2021 | QuCloud: A New Qubit Mapping Mechanism for Multi-programming Quantum Computing in Cloud EnvironmentabstractFor a specific quantum chip, multi-programming improves overall throughput and resource utilization. Previous studies on mapping multiple programs often lead to resource under-utilization, high error rate, and low fidelity. This paper proposes QuCloud, a new approach for mapping quantum programs in the cloud environment. We have three new designs in QuCloud. (1) We leverage the community detection technique to partition physical qubits among concurrent quantum programs, avoiding the waste of robust resources. (2) We design X-SWAP scheme that enables inter-program SWAPs and prioritizes SWAPs associated with critical gates to reduce the SWAP overheads. (3) We propose a compilation task scheduler that schedules concurrent quantum programs to be compiled and executed based on estimated fidelity for the best practice. We evaluate our work on publicly available quantum computer IBMQ16 and a simulated quantum chip IBMQ50. Our work outperforms the state-of-the-art work for multi-programming on fidelity and compilation overheads by 9.7% and 11.6%, respectively. Lei Liu 0037, Xinglei Dou |
HPCA | 1 |
| 2020 | A New Qubits Mapping Mechanism for Multi-programming Quantum ComputingabstractFor a specific quantum chip, multi-programming helps to improve the overall throughput and resource utilization. However, previous solutions for mapping multiple programs often lead to resource under-utilization, high error rate, and low fidelity. In this paper, we propose a new approach to map concurrent quantum programs. Our approach has three critical components. The first one is the Community Detection Assisted Partition (CDAP) algorithm, which partitions physical qubits for concurrent quantum programs by considering physical typology and the error rate, avoiding the waste of robust resources. The second one is the X-SWAP scheme that enables inter-program SWAPs and prioritizes SWAPs associated with critical gates to reduce the SWAP overheads. Finally, we propose a compilation task scheduler, which dynamically selects concurrent quantum programs to be compiled and executed together based on estimated fidelity for the best practice. We evaluate our work on publicly available quantum computer IBMQ16 and a simulated quantum chip IBMQ50. Our work outperforms the state-of-the-art work for multi-programming on fidelity and compilation overheads by 12.0% and 11.6%, respectively. Xinglei Dou, Lei Liu 0037 |
PACT | 2 |
| 2020 | Monitoring Memory Behaviors and Mitigating NUMA Drawbacks on Tiered NVM Systems
Shengjie Yang, Xinglei Dou, Xiaoli Gong, Hao Liu 0107, Lei Liu 0037 |
NPC | 7 |
| 2020 | Architectural Support for NVRAM Persistence in GPUsabstractNon-volatile Random Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices, such as Solid State Drives (SSD). In addition to higher storage density, NVRAM provides byte-addressability, higher bandwidth, near-DRAM latency, and easier access compared to block devices such as traditional SSDs. This enables new programming paradigms taking advantage of durability and larger memory footprint. With the range and size of GPU workloads expanding, NVRAM will present itself as a promising addition to GPU's memory hierarchy. To utilize the non-volatility of NVRAMs, programs should allow durable stores, maintaining consistency through a power loss event. This is usually done through a logging mechanism that works in tandem with a transaction execution layer which can consist of a transactional memory or a locking mechanism. Together, this results in a transaction processing system that preserves the ACID properties. GPUs are designed with high throughput in mind, leveraging high degrees of parallelism. Transactional memory proposals enable fine-grained transactions at the GPU thread-level. However, with lower write bandwidths compared to that of DRAMs, using NVRAM as-is may yield sub-optimal overall system performance when threads experience long latency. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases. Due to the speedup, our proposed method also results in reduction in overall energy consumption. Sui Chen, Lei Liu 0037, Lu Peng 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Efficient GPU NVRAM Persistence with Helper WarpsabstractNon-volatile Random-Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices. To utilize the non-volatility of NVRAMs, programs should allow durable stores, meaning consistency must be maintained during a power loss event. GPUs are designed with high throughput, leveraging high degrees of parallelism. However, with lower NVRAM write bandwidths compared to that of DRAMs, using NVRAM as is may yield suboptimal overall system performance. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases. Sui Chen, Faen Zhang, Lei Liu 0037, Lu Peng 0001 |
DAC | 3 |
| 2019 | Hierarchical Hybrid Memory Management in OS for Tiered Memory SystemsabstractThe emerging hybrid DRAM-NVM architecture is challenging the existing memory management mechanism at the level of the architecture and operating system. In this paper, we introduce Memos, a memory management framework which can hierarchically schedule memory resources over the entire memory hierarchy including cache, channels, and main memory comprising DRAM and NVM simultaneously. Powered by our newly designed kernel-level monitoring module that samples the memory patterns by combining TLB monitoring with page walks, and page migration engine, Memos can dynamically optimize the data placement in the memory hierarchy in response to the memory access pattern, current resource utilization, and memory medium features. Our experimental results show that Memos can achieve high memory utilization, improving system throughput by around 20.0 percent; reduce the memory energy consumption by up to 82.5 percent; and improve the NVM lifetime by up to 34X. Lei Liu 0037, Shengjie Yang, Lu Peng 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |