VLDB 2026 Research / reviewers in the wild / expert
Zhinan Cheng
dblp:160/2167
· DBLP profile ↗
9ranked-venue papers
6as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pulse: Fine-Grained Hierarchical Hashing Index for Disaggregated MemoryabstractBy decoupling compute and memory resources into independent pools that are provisioned and managed separately, disaggregated memory (DM) is promising to break the scaling constraints for memory systems and improve resource utilization. However, it also comes with a new challenge to design a high-performance hashing index to manage the vast memory pool with weak computing power. In this paper, we reconsider this problem and find that existing hashing indexes for DM still experience two fundamental yet unresolved limitations: (i) amplifying traffic under high insertion concurrency, and (ii) introducing significantly high insertion latency, stemming mainly from the directory synchronization and item relocation in resizing process. We resolve the above limitations by designing Pulse, a finegrained hierarchical hashing index for DM. Pulse comprises the following three design primitives. It proposes a multi-level index structure, which breaks the conventional flat directory into multiple sub-directories that are organized hierarchically, achieving fine-grained directory synchronization. Pulse also maintains a small portion of hashed keys in the memory pool, which aids in item relocation during resizing, thereby reducing resizing traffic and promising system stability. Pulse finally exploits operation parallelism by tailoring the doorbell batching mechanism with selective signaling. We conduct extensive experiments using a variety of benchmarks, showing that Pulse can improve$3.46 \times$of the throughput and reduce 76.3% of the tail latency compared to state-of-the-art hashing indexes. Guangyang Deng, Zixiang Yu, Zhirong Shen, Qiangsheng Su, Zhinan Cheng, Jiwu Shu |
HPCA | 5 |
| 2022 | An In-Depth Correlative Study Between DRAM Errors and Server Failures in Production Data CentersabstractDynamic Random Access Memory (DRAM) errors are prevalent and lead to server failures in production data centers. However, little is known about the correlation between DRAM errors and server failures in state-of-the-art field studies on DRAM error measurement. To fill this void, we present an in-depth data-driven correlative analysis between DRAM errors and server failures, with the primary goal of predicting server failures based on DRAM error characterization and hence enabling proactive reliability maintenance for production data centers. Our analysis is based on an eight-month dataset collected from over three million memory modules in the production data centers at Alibaba. We find that the correctable DRAM errors of most server failures only manifest within a short time before the failures happen, implying that server failure prediction should be conducted regularly at short time intervals for accurate prediction. We also study various impacting factors (including component failures in the memory subsystem, DRAM configurations, types of correctable DRAM errors) on server failures. Furthermore, we design a machine-learning-based server failure prediction workflow and demonstrate the feasibility of server failure prediction based on DRAM error characterization. To this end, we report 14 findings from our measurement and prediction studies. Zhinan Cheng, Shujie Han 0001, Patrick P. C. Lee, Jiongzhou Liu |
SRDS | 1 |
| 2021 | Enabling Low-Redundancy Proactive Fault Tolerance for Stream Machine Learning via Erasure CodingabstractMachine learning for continuous data streams, or stream machine learning in short, is increasingly adopted in real-time big data applications. Fault tolerance is a critical requirement for stream machine learning applications in large-scale distributed deployment. However, existing reactive fault tolerance mechanisms, which trigger failure recovery upon the detection of failures, inevitably incur high recovery overhead and compromise the low-latency requirement of stream machine learning. We design StreamLEC, a stream machine learning system that leverages erasure coding to provide low-redundancy proactive fault tolerance for immediate failure recovery. StreamLEC supports general stream machine learning applications, and incorporates different techniques to mitigate erasure coding overhead. Evaluation on a local cluster and Amazon EC2 shows that StreamLEC achieves much higher throughput than both reactive fault tolerance and replication-based proactive fault tolerance, with negligible failure recovery overhead. Zhinan Cheng, Lu Tang 0004, Qun Huang 0001, Patrick P. C. Lee |
SRDS | 1 |
| 2021 | Automated Intelligent Healing in Cloud-Scale Data CentersabstractModern cloud-scale data centers necessitate self-healing (i.e., the automation of detecting and repairing component failures) to support reliable and scalable cloud services in the face of prevalent failures. Traditional policy-based self-healing solutions rely on expert knowledge to define the proper policies for choosing repair actions, and hence are error-prone and non-scalable in practical deployment. We propose AIHS, an automated intelligent healing system that applies machine learning to achieve scalable self-healing in cloud-scale data centers. AIHS is designed as a full-fledged, general pipeline that supports various machine learning models for predicting accurate repair actions based on raw monitoring logs. We conduct extensive trace-driven and production experiments, and show that AIHS achieves higher prediction accuracy than current self-healing solutions and successfully fixes 92.4% of the total of 33.7 million production failures over seven months. AIHS also reduces 51% of unavailable time of each failed server on average compared to policy-based self-healing. AIHS is now deployed in production cloud-scale data centers at Alibaba with a total of 600 K servers. We open-source a Python prototype that reproduces the self-healing pipeline of AIHS for public validation. Zhinan Cheng, Patrick P. C. Lee, Pinghui Wang, Yi Qiang, Jinlong Lu, Xinquan Ding |
SRDS | 2 |
| 2019 | On the performance and convergence of distributed stream processing via approximate fault tolerance
Zhinan Cheng, Qun Huang 0001, Patrick P. C. Lee |
VLDB J. | 1 |
| 2016 | Display power reduction for mobile closed-source gamesabstractWith the rapid development of mobile games, power consumption and battery life of mobile platforms become a crucial problem, therefore, saving power for display which consumes a lot of power without decreasing the user experience becomes necessary. Content-centric techniques are most commonly used to save display power. However, they cannot be applied to closed-source games as they need to modify the game images. In this paper, we explore a backlight power model and propose a novel method to make a trade-off between the game image quality and display power. Different from content-centric methods, based on the proposed trade-off model, we propose a backlight dimming algorithm which maintains the user game experience using game-state information and saves display power for mobile games without modifying any game image. We implement the proposed trade-off model and backlight dimming policy in a contemporary mobile platform where the evaluation results show that, maintaining a specific game image quality level, our policy can save system power up to 10.43% compared with the static policy without decreasing games' performances. Zhinan Cheng, Xi Li 0003, Jiachen Song, Beilei Sun, Xuehai Zhou, Chao Wang 0003 |
ASAP | 1 |
| 2016 | FCM: Towards Fine-Grained GPU Power Management for Closed Source Mobile GamesabstractContemporary mobile platforms employ embedded graphic processing units (GPUs) for graphics-intensive games, and dynamic voltage and frequency scaling (DVFS) policies are used to save energy without sacrificing quality. However, current GPU DVFS policies result in unnecessary power waste due to defective workload estimations of embedded GPUs during game play. In this paper, we propose the Frame-Complexity Model (FCM), a fine-grained estimation of the GPU workload in a game frame, to quantify the GPU workload with the real runtime demand for GPU computing resources of a game frame. In FCM, three constituents of a game frame (i.e., structure, textures and computation) are quantified without modification of mobile games. Preliminary experiments show that, compared with the default policy, the FCM-directed GPU DVFS policy can reduce more power consumption of games (11.3% to 25.8%) with good Quality of Service (QoS). Jiachen Song, Xi Li 0003, Beilei Sun, Zhinan Cheng, Chao Wang 0003, Xuehai Zhou |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Behavior-Aware Integrated CPU-GPU Power Management for Mobile GamesabstractSince game applications have spilled over on the modern mobile platforms equipped with Multiprocessor Systemon-Chips and highlighted the power consumption and battery life problem of these platforms, reducing the game power for mobile devices becomes meaningful. The design of independent CPU-GPU power managements in contemporary platforms results in power consumption waste due to the failure of consideration of CPU-GPU interaction and game workload behaviors. Through analyzing the Application-Operating System (APP-OS) interaction and CPU-GPU interaction, we extract the system-call information and OpenGL API information to characterize the game workload in a low-complexity way. In this paper, based on identifying the game workload behavior and performance bottleneck, we propose a behavior-aware integrated CPU-GPU power management approach for mobile games. We also implement our power saving policy in the real platform, where the evaluation results show that our behavior-aware policy can significantly reduce power and improve game performance. Our policy provides 18% and 5% higher power-efficiency on average compared with the current policy used in our platform and the state-of-the-art policy respectively. Zhinan Cheng, Xi Li 0003, Beilei Sun, Jiachen Song, Chao Wang 0003, Xuehai Zhou |
MASCOTS | 1 |
| 2015 | Automatic frame rate-based DVFS of gameabstractThe rapid development of mobile games highlights the power consumption problem in the mobile platform. Most of the power saving techniques use the prediction-based dynamic voltage frequency scaling (DVFS) scheme. However, the prediction could be inaccurate resulting from the frequent interactions of user when playing games. We have observed that frame rate is near-linear to CPU frequency, but there is a bottleneck, frame rate will not increase as CPU frequency increases when CPU frequency reaches this threshold. Moreover, previous research has shown that utilizing the information of game state can reduce the influence of game interactive characterization to DVFS policy. We explore a method to automatically detect the game state. We propose the Automatic Frame Rate-Based DVFS policy, which can learn the threshold of frame rate online and utilize the information of game state and frame rate to scale the frequency without prediction. Our evaluation result shows that, compared with the prediction-based Android default Interactive DVFS policy, our policy saves more power in all the testing games. Up to 15.2% more power can be saved by Automatic Frame Rate-Based DVFS policy. Zhinan Cheng, Xi Li 0003, Beilei Sun, Ce Gao, Jiachen Song |
ASAP | 1 |