Zhihai Huang

dblp:312/9912 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021
YearPublicationVenuePosition
2026 Mitigating CDN Cache Misses with Scheduling: An Origin Shield for Billion-QPS Social Platforms
abstract
The explosive growth of modern short-video platforms like WeChat Channels — approaching 1.4 billion monthly active users (MAU) while serving over one billion queries per second (QPS) — has exposed fundamental limitations in conventional content delivery network (CDN) architectures, particularly their inability to handle highly dynamic content popularity patterns. The traditional cache-centric approach proves to be catastrophic when faced with ephemeral content exhibiting a "million-item waist" distribution. This phenomenon results in severe cache pollution (manifesting as 5–40% origin miss ratios) and introduces prohibitive back-to-origin bandwidth costs. Our TritonFlow overcomes these challenges by introducing access point scheduling as a novel origin traffic compensator, which reduces the operational cost of back-to-origin fetches by enhancing CDN cache affinity. Through analysis of production traces (peak >1B QPS), we demonstrate that TritonFlow reduces scheduling fluctuations by 96% and decreases back-to-origin traffic caused by non-compulsory cache misses by 35%, all while maintaining sub-200ms fetch latency.
Zixuan Yang 0003, Jiaqi Zheng 0001, Boxi Liu, Guihai Chen, Quan Xia, Zhihai Huang, Shangce Yuan
EuroSys8
2024 CGHit: A Content-Oriented Generative-Hit Framework for Content Delivery Networks
abstract
The service provided by content delivery networks (CDNs) may overlook content locality, leaving the potential to improve performance. In this study, we explore the feasibility of leveraging generated data as a replacement for fetching data in missing scenarios based on content locality. Due to sufficient local computing resources and reliable generation efficiency, we propose a content-oriented generative-hit framework (CGHit) for CDNs. CGHit utilizes idle computing resources on edge nodes to generate requested data based on similar or related cached data, achieving hits. Extensive experiments in a real-world system demonstrate that CGHit reduces the average access latency by half. In addition, experiments conducted on a simulator confirm that CGHit can enhance current caching algorithms, leading to lower latency and reduced bandwidth usage.
Peng Wang 0037, Yu Liu 0040, Ke Liu 0014, Ke Zhou 0001, Zhihai Huang
NAS8
2024 $\varepsilon$ɛ-LAP: A Lightweight and Adaptive Cache Partitioning Scheme With Prudent Resizing Decisions for Content Delivery Networks
abstract
As dependence on Content Delivery Networks (CDNs) increases, there is a growing need for innovative solutions to optimize cache performance amid increasing traffic and complicated cache-sharing workloads. Allocating exclusive resources to applications in CDNs boosts the overall cache hit ratio (OHR), enhancing efficiency. However, the traditional method of creating the miss ratio curve (MRC) is unsuitable for CDNs due to the diverse sizes of items and the vast number of applications, leading to high computational overhead and performance inconsistency. To tackle this issue, we propose alightweight andadaptive cachepartitioning scheme called$\varepsilon$-LAP. This scheme uses a corresponding shadow cache for each partition and sorts them based on the average hit numbers on the granularity unit in the shadow caches. During partition resizing,$\varepsilon$-LAP transfers storage capacity, measured in units of granularity, from the$(N-k+1)$-th ($k\leq \frac{N}{2}$) partition to the$k$-th partition. A learning threshold parameter, i.e.,$\varepsilon$, is also introduced to prudently determine when to resize partitions, improving caching efficiency. This can eliminate about 96.8% of unnecessary partition resizing without compromising performance.$\varepsilon$-LAP, when deployed inPicCloudatTencent, improved OHR by 9.34% and reduced the average user access latency by 12.5 ms. Experimental results show that$\varepsilon$-LAP outperforms other cache partitioning schemes in terms of both OHR and access latency, and it effectively adapts to workload variations.
Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Liu 0014, Ke Zhou 0001, Zhihai Huang
IEEE Trans. Cloud Comput.7
2024 Beyond Belady to Attain a Seemingly Unattainable Byte Miss Ratio for Content Delivery Networks
abstract
Reducing the byte miss ratio (BMR) in the Content Delivery Network (CDN) caches can help providers save on the cost of paying for traffic. When evicting objects or files of different sizes in the caches of CDNs, it is no longer sufficient to pursue an optimal object miss ratio (OMR) by approximating Belady to ensure an optimal BMR. Our experimental observations suggest that there are multiple request sequence windows. In these windows, a replacement policy prioritizes the eviction of objects with large sizes and ultimately evicts the object with the longest reuse distance, lowering the BMR without increasing the OMR. To accurately capture those windows, we monitor the changes in OMR and BMR using a deep reinforcement learning (RL) model and then implement a BMR-friendly replacement algorithm in these windows. Based on this policy, we propose a Belady and Size Eviction (LRU-BaSE) algorithm that reduces BMR while maintaining OMR. To make LRU-BaSE efficient and practical, we address the feedback delay problem of RL with a two-pronged approach. On the one hand, we shorten the LRU-base decision region based on the observation that the rear section of the cache queue contains most of the eviction candidates. On the other hand, the request distribution on CDNs makes it feasible to divide the learning region into multiple sub-regions that are each learned with reduced time and increased accuracy. In real CDN systems, LRU-BaSE outperforms LRU by reducing “backing to OS” traffic and access latency by 30.05% and 17.07%, respectively, on average. In simulator tests, LRU-BaSE outperforms state-of-the-art cache replacement policies. On average, LRU-BaSE's BMR is 0.63% and 0.33% less than that of Belady and Practical Flow-based Offline Optimal (PFOO), respectively. In addition, compared to Learning Relaxed Belady (LRB), LRU-BaSE can yield relatively stable performance when facing workload drift.
Peng Wang 0037, Hong Jiang 0001, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang
IEEE Trans. Parallel Distributed Syst.6
2023 Smart Cache Insertion and Promotion Policy for Content Delivery Networks
abstract
Improving hit rates can be achieved by enhancing cache replacement algorithms with the identification of zero-reuse objects (ZROs) and inserting them at the end of the cache queue. Note that the promotion policy needs to achieve a similar task as the above insertion policy since the hit object may immediately become a ZRO (called P-ZRO) that is not suitable for placement at the front of the queue. However, existing studies have yet to consider P-ZROs, and current insertion algorithms struggle to simultaneously identify both ZROs and P-ZROs. To address these issues, we propose integrating the insertion and promotion policies. We do this by treating hit objects as special missing objects and employing reinforcement learning to create a unified model for both policies, where the learning function recognizes the relationship between performance changes and the emergence of ZROs and P-ZROs. Our proposed solution is a smart cache insertion and promotion policy (SCIP) that dynamically adjusts the insertion position using a bimodal insertion policy for both missing and hit objects, guided by the model. Extensive experiments demonstrate that SCIP significantly improves overall performance in real-world content delivery network systems and outperforms state-of-the-art insertion policies in terms of miss ratios in the simulator. In addition, deploying SCIP on optimal cache replacement algorithms can further decrease their miss ratios.
Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen
ICPP5
2022 Adaptive Size-Aware Cache Insertion Policy for Content Delivery Networks
abstract
Content delivery networks (CDNs) are large distributed cache systems that deliver objects with inconsistent sizes. The zero-reuse objects that are not reused in a time window but still loaded and evicted in the cache waste cache resources and result in degradation of object hit ratio (OHR) in CDNs. Although prohibiting these objects from entering the cache is a viable solution, the variable workloads and various object sizes in CDNs make the determination of zero-reuse difficult, resulting in an increased risk of bandwidth overhead in the data center by the misjudgment. To alleviate this problem, we propose to use the insertion policy to give each object at least one chance to be hit. Meanwhile, we find that the distribution of zero-reuse objects correlates with their sizes through data analysis. As a result, we propose an adaptive size-aware cache insertion policy (ASC-IP) for the OHR improvement and design an adaptive scheme to dynamically adjust the size threshold used to determine the zero-reuse objects, adapting the mutative access patterns with negligible overhead. We have deployed ASC-IP in TDC of Company-T and ASC-IP can improve the OHR by 9.6% and reduce the user access latency by 7.14ms on average and reduce the back-to-source bandwidth by 8.75Gbps. In addition, on Twitter, Wikipedia, and a real-world Trace-T, we show that ASC-IP outperforms state-of-the-art cache algorithms working on CDNs and can upgrade LRU-based replacement algorithms with negligible overheads.
Peng Wang 0037, Yu Liu 0040, Zhelong Zhao, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen
ICCD5
2022 A Lightweight and Adaptive Cache Partitioning Scheme for Content Delivery Networks
abstract
Allocating exclusive resources for different applications in content delivery networks (CDNs) allows for a higher overall hit ratio. The cache partitioning schemes on Last-Level Cache (LLC) are promising solutions that dynamically split cache sizes into partitions corresponding to threads by the miss ratio curve (MRC). Nonetheless, due to the sheer number of applications and various item sizes in CDNs, partitioning via MRC will cause high computational overheads and performance fluctuations. As a result, in this paper, we propose a lightweight and adaptive cache partitioning scheme (LAP) for CDNs. LAP establishes a shadow cache for each partition, where the size of the partition and its shadow cache is equal to the size of the integral cache. The average number of hits on the granularity unit in the shadow caches, where the size of the granularity equals the size of the probable largest item, is used to sort N partitions in decreasing order. When resizing partitions, LAP transfers a capacity of the size of granularity from the (N – k + 1)-th $\left( {k \leq \frac{N}{2}} \right)$ partition into the k-th partition. Meanwhile,we provide a threshold that neglects partition resizing and improves partitioning efficiency. This lightweight scheme can enhance resource utilization by progressively adapting to workload variations. We have deployed LAP in PicCloud of Company-T and LAP can improve the OHR by 9.34% and reduce the average user access latency by 12.5ms. Then, we verify LAP in the public trace from Akamai and the real trace from PicCloud. Experimental results demonstrate that LAP outperforms other cache partitioning schemes and tackles the performance cliff problem with little overhead.
Peng Wang 0037, Zhelong Zhao, Yu Liu 0040, Ke Zhou 0001, Zhihai Huang, Yanxiong Chen
ICCD5