VLDB 2026 Research / reviewers in the wild / expert
Zheng Liu 0022
dblp:06/3580-22
· DBLP profile ↗
19ranked-venue papers
1as first author
18since 2021 · last 2026
0009-0001-6688-4115ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 8 since 2021Software engineering, systems software and programming languages · 7 · 7 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nebula: Infinite-Scale 3D Gaussian Splatting in VR via Collaborative Rendering and Accelerated Stereo Rasterizationabstract3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, current architectural designs often overlook the 3DGS scalability, making them fragile for extremely large-scale 3DGS. Meanwhile, the VR bandwidth requirement makes it impossible to deliver high-fidelity and smooth VR content from the cloud. Zheng Liu 0022, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, Yu Feng 0007 |
ASPLOS (2) | 2 |
| 2026 | LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent OptimizationabstractData streaming operations consume a significant portion of CPU resources in data centers. The Data Streaming Accelerator (DSA), integrated into modern Intel CPUs in datacenter, offers promising acceleration for these operations with user-friendly features. However, previous studies have overlooked DSA's internal mechanisms and the performance implications of these features, leaving key performance issues unresolved in real-world usage. Yuansen Wang, Teng Ma 0006, Yuanhui Luo, Dongbiao He, Zheng Liu 0022, Yunpeng Chai |
EuroSys | 5 |
| 2026 | PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
Chen Gong 0005, Zheng Liu 0022, Kecen Li, Tianhao Wang 0001 |
NDSS | 2 |
| 2026 | PrivCode: When Code Generation Meets Differential Privacy
Zheng Liu 0022, Chen Gong 0005, Terry Yue Zhuo, Kecen Li, Weichen Yu, Matt Fredrikson, Tianhao Wang 0001 |
NDSS | 1 |
| 2025 | StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic TerminationabstractPoint clouds are increasingly important in intelligent applications, but frequent off-chip memory traffic in accelerators causes pipeline stalls and leads to high energy consumption. While conventional line buffer techniques can eliminate off-chip traffic, they cannot be directly applied to point clouds due to their inherent computation patterns. To address this, we introduce two techniques: compulsory splitting and deterministic termination, enabling fully-streaming processing. We further propose StreamGrid, a framework that integrates these techniques and automatically optimizes on-chip buffer sizes. Our evaluation shows StreamGrid reduces on-chip memory by 61.3% and energy consumption by 40.5% with marginal accuracy loss compared to the baselines without our techniques. Additionally, we achieve 10.0× speedup and 3.9× energy efficiency over state-of-the-art accelerators. Yu Feng 0007, Zheng Liu 0022, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Zhezhi He, Jieru Zhao, Yuhao Zhu 0001 |
ASPLOS (2) | 2 |
| 2025 | CXL-INTERPLAY: Unraveling and Characterizing CXL Interference in Modern Computer SystemsabstractCompute Express Link (CXL) is a promising technology that addresses memory and storage challenges. Despite its advantages, CXL faces performance threats from external interference when coexisting with current memory and storage systems. This interference is under-explored in existing research. To address this, we develop CXL-Interplay, systematically characterizing and analyzing interference from memory and storage systems. To the best of our knowledge, we are the first to characterize CXL interference on real CXL hardware. We also provide reverse-reasoning analysis with performance counters and kernel functions. In the end, we propose and evaluate mitigation solutions. Shunyu Mao, Jiajun Luo, Jiapeng Zhou, Zheng Liu 0022, Teng Ma 0006, Shuwen Deng |
DAC | 6 |
| 2025 | DSA-2LM: A CPU-Free Tiered Memory Architecture with Intel DSA
Ruili Liu, Teng Ma 0006, Yingdi Shan, Zheng Liu 0022, Lingfeng Xiang, Hui Lu 0001, Jia Rao, Kang Chen 0001, Yongwei Wu 0001 |
USENIX ATC | 6 |
| 2025 | MemTunnel: A CXL-Based Rack-Scale Host Memory Pooling Architecture for Cloud ServiceabstractMemory underutilization poses a significant challenge in cloud services, leading to performance inefficiencies and resource wastage. The tightly coupled computing and memory resources in cloud servers are identified as the root cause of this problem. To address this issue, memory pooling has been the subject of extensive research for decades, providing centralized or distributed shared memory pools as flexible memory resources for various applications running on different servers. However, existing memory disaggregation solutions sacrifice memory resources, add extra hardware (such as memory boxes/blades/drives), and degrade memory performance to achieve flexibility. To overcome these limitations, this paper proposes MemTunnel, a rack-scale host memory pooling architecture that provides a low-cost memory pooling solution based on Compute Express Link (CXL). MemTunnel is the first hardware and software architecture to offer symmetric, memory-semantic memory pooling over CXL, with an FPGA-based platform to demonstrate its feasibility in a real implementation. MemTunnel is orthogonal to the existing CXL-based memory pool and provides an additional layer of abstraction for memory disaggregation. Evaluation results show that MemTunnel achieves comparable performance to the existing CXL-based memory pool for a single machine and provides better rack-scale performance with minor hardware overheads. Tianchan Guan, Yijin Guan, Zhaoyang Du, Jiacheng Ma 0001, Boyu Tian, Teng Ma 0006, Zheng Liu 0022, Yuan Xie 0001, Mingyu Gao 0001, Guangyu Sun 0003, Hongzhong Zheng, Dimin Niu |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2025 | LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based AnalysisabstractLog-based AIOps is a widely researched topic aiming at reducing the developer burden in system maintenance. Since industrial developers prefer lightweight supervised solutions for log-based AIOps, the strong dependence of these solutions on labeled data creates significant challenges for teams new to building log-based AIOps capabilities, such as high labeling costs, inconsistent annotations, and manual management issues. Log-labeling faces challenges in integrating existing artifacts to reduce labeling costs and manage labels effectively. To the best of our knowledge, no prior research addresses of assisting log-labeling problem. In this article, we propose a new approach called LogLabeler to assist developers in annotating and managing log labels. LogLabeler leverages existing artifacts for initial-label-acquisition, minimizes labeling costs by automatically generating all log labels, and shields developers from manual label management through a human-in-the-loop refinement approach. Evaluations on real-world datasets from Alibaba and open-source datasets show that LogLabeler can effectively supplement log labels, achieving comparable accuracy to existing baselines while operating more efficiently. Furthermore, we demonstrate LogLabeler's practical effectiveness at Alibaba through a case study, highlighting its benefits to developers. Zongyang Li, Qinglong Wang 0003, Shangming Cai, Zheng Liu 0022, Tao Ma 0006, Wei Yang 0013, Ying Li 0012, Tao Xie 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and NodesabstractServerless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform with the operating system and CXL/RDMA-based remote memory pools in two key areas. Firstly, TrEnv introduces repurposable sandboxes, which can be shared across different functions and hence, substantially decrease the overhead associated with creating isolation sandboxes. Secondly, it augments the OS with "memory templates" that enable rapid restoration of function states stored on remote memory. These innovations allow TrEnv to facilitate rapid transitions between instances of different functions and enable memory sharing across multiple nodes. Our evaluations using a variety of representative and real-world workloads demonstrate that TrEnv can initiate a container within 10 milliseconds, achieving up to a 7× speedup in P99 end-to-end latency and reducing memory usage by 48% on average compared to state-of-the-art on-demand restoring systems. Teng Ma 0006, Zheng Liu 0022, Sixing Lin, Kang Chen 0001, Jinlei Jiang, Xia Liao, Yingdi Shan, Mengting Lu, Tao Ma 0006, Haifeng Gong, Yongwei Wu 0001 |
SOSP | 4 |
| 2024 | Diagnosing Application-network Anomalies for Millions of IPs in Production Clouds
Zhe Wang 0015, Huanwu Hu, Linghe Kong, Xinlei Kang, Qiao Xiang, Peihao Yang, Jiejian Wu, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Xianlong Zeng, Dennis Cai, Guihai Chen |
USENIX ATC | 13 |
| 2024 | HydraRPC: RPC in the CXL Era
Teng Ma 0006, Zheng Liu 0022, Chengkun Wei, Youwei Zhuo, Yijin Guan, Dimin Niu, Tao Ma 0006 |
USENIX ATC | 2 |
| 2024 | Zero+: Monitoring Large-Scale Cloud-Native Infrastructure Using One-Sided RDMAabstractCloud services have shifted from monolithic designs to microservices running on cloud-native infrastructure with monitoring systems to ensure service level agreements (SLAs). However, traditional monitoring systems no longer meet the demands of cloud-native monitoring. In Alibaba’s “double eleven” shopping festival, it is observed that the monitor occupies resources of the monitored infrastructure and even disrupts services. In this paper, we propose a novel monitoring system named for cloud-native monitoring. achieves zero overhead in collecting raw metrics using one-sided remote direct memory access (RDMA) and remedies network congestion by adopting a receiver-driven flow control scheme. also features a priority queue mechanism to meet different quality of service requirements and an efficient batch processing design to relieve CPU occupation. has been deployed and evaluated in four different clusters with heterogeneous RDMA NIC devices and architectures in Alibaba Cloud. Results show that achieves no CPU occupation at the monitored host and supports$1\sim10k$hosts with$0.1\sim1s$sampling interval using a single thread for network I/O. significantly relieves the incast issue and maintains$80\sim95\%$of bandwidth utilization in several clusters when monitoring$1k$hosts. also ensures services with high priority accomplish collecting metrics earlier than low priority ones by at least$400 \mu s$when monitoring$1k$hosts. Jiejian Wu, Teng Ma 0006, Zhe Wang 0015, Linghe Kong, Zhenzao Wen, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Guihai Chen |
IEEE/ACM Trans. Netw. | 11 |
| 2023 | LigBee: Symbol-Level Cross-Technology Communication from LoRa to ZigBeeabstractLow-power wide-area networks (LPWAN) evolve rapidly with advanced communication primitives (e.g., coding, modulation) being continuously invented. This rapid iteration on LPWAN, however, forms a communication barrier between legacy wireless sensor nodes deployed years ago (e.g., ZigBee-based sensor node) with their latest competitor running a different communication protocol (e.g., LoRa-based IoT node): they work on the same frequency band but share different MAC- and PHY-layer regulations and thus cannot talk to each other directly. To break this barrier, we propose LigBee, a cross-technology communication (CTC) solution that enables symbol-level communication from the latest LPWAN LoRa node to legacy ZIGBEE node. We have implemented LigBee on both software-defined radios and commercial-off-the-shelf (COTS) LoRa and ZigBee nodes, and demonstrated that LigBee builds a reliable CTC link from LoRa node to ZigBee node on both platforms. Our experimental results show that i) LigBee achieves a bit error rate (BER) in the order of 10−3with 70 ∼ 80% frame reception ratio (FRR), ii) the range of LigBee link is over 300m, which is 6 ∼ 7.5× the typical range of legacy ZigBee and state-of-the-art solution, and iii) the throughput of LigBee link is maintained on the order of kbps, which is close to the LoRa’s throughput. Zhe Wang 0015, Linghe Kong, Longfei Shangguan, Liang He 0002, Kangjie Xu, Yifeng Cao, Qiao Xiang, Jiadi Yu, Teng Ma 0006, Zheng Liu 0022, Guihai Chen |
INFOCOM | 12 |
| 2023 | KeenTune: Automated Tuning Tool for Cloud Application Performance Testing and OptimizationabstractThe performance testing and optimization of cloud applications is challenging, because manual tuning of cloud computing stacks is tedious and automated tuning tools are rare used for cloud services. To address this issue, we introduce KeenTune, an automated tuning tool designed to optimize application performance and facilitate performance testing. KeenTune is a lightweight and flexible tool that can be deployed with to-be-tuned applications with negligible impact on their performance. Specifically, KeenTune uses a surrogate model that can be implemented with machine learning models to filter out less relevant parameters for efficient tuning. Our empirical evaluation shows that KeenTune significantly enhances the throughput performance of Nginx web servers, resulting in performance improvements of up to 90.43% and 117.23% in certain cases. This study highlights the benefits of using KeenTune for achieving efficient and effective performance testing of cloud applications. The video and source code for KeenTune are provided as supplementary materials. Qinglong Wang 0003, Runzhe Wang, Xiaohai Shi, Zheng Liu 0022, Tao Ma 0006, Houbing Song, Heyuan Shi |
ISSTA | 5 |
| 2023 | Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared MemoryabstractThe efficiency of distributed shared memory (DSM) has been greatly improved by recent hardware technologies. But, the difficulty of distributed memory management can still be a major obstacle to the democratization of DSM, especially when a partial failure of the participating clients (e.g., due to crashed processes or machines) should be tolerated. Teng Ma 0006, Jinqi Hua, Zheng Liu 0022, Kang Chen 0001, Fan Du, Jinlei Jiang, Tao Ma 0006, Yongwei Wu 0001 |
SOSP | 4 |
| 2023 | Async-fork: Mitigating Query Latency Spikes Incurred by the Fork-based Snapshot Mechanism from the OS LevelabstractIn-memory key-value stores (IMKVSes) serve many online applications. They generally adopt the fork-based snapshot mechanism to support data backup. However, this method can result in query latency spikes because the engine is out-of-service for queries during the snapshot. In contrast to existing research optimizing snapshot algorithms, we address the problem from the operating system (OS) level, while keeping the data persistent mechanism in IMKVSes unchanged. Specifically, we first study the impact of the fork operation on query latency. Based on findings in the study, we propose Async-fork, which performs the fork operation asynchronously to reduce the out-of-service time of the engine. Async-fork is implemented in the Linux kernel and deployed into the online Redis database in public clouds. Our experiment results show that Async-fork can significantly reduce the tail latency of queries during the snapshot. Pu Pang, Kaihao Bai, Quan Chen 0002, Shixuan Sun, Bo Liu 0122, Hongbo Yao, Zhengheng Wang, Zheng Liu 0022, Yong Yang 0013, Tao Ma 0006, Minyi Guo |
Proc. VLDB Endow. | 11 |
| 2022 | Industry practice of configuration auto-tuning for cloud applications and servicesabstractAuto-tuning attracts increasing attention in industry practice to optimize the performance of a system with many configurable parameters. It is particularly useful for cloud applications and services since they have complex system hierarchies and intricate knob correlations. However, existing tools and algorithms rarely consider practical problems such as workload pressure control, the support for distributed deployment, and expensive time costs, etc., which are utterly important for enterprise cloud applications and services. In this work, we significantly extend an open source tuning tool – KeenTune to optimize several typical enterprise cloud applications and services. Our practice is in collaboration with enterprise users and tuning tool developers to address the aforementioned problems. Specifically, we highlight five key challenges from our experiences and provide a set of solutions accordingly. Through applying the improved tuning tool to different application scenarios, we achieve 2%-14% improvements for the performance of MySQL, OceanBase, nginx, ingress-nginx, and 5%-70% improvements for the performance of ACK cloud container service. Runzhe Wang, Qinglong Wang 0003, Heyuan Shi, Yuheng Shen, Zheng Liu 0022, Xiaohai Shi, Yu Jiang 0001 |
ESEC/SIGSOFT FSE | 8 |
| 2020 | Spool: Reliable Virtualized NVMe Storage Pool in Public Cloud Infrastructure
Shang Zhao 0003, Quan Chen 0002, Zheng Liu 0022, Tao Ma 0006, Yong Yang 0013, Yanbo Zhou, Keqiang Niu, Sijie Sun, Minyi Guo |
USENIX ATC | 5 |