Yazhuo Zhang

dblp:22/8903 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Rethinking Web Cache Design for the AI Era
abstract
Web caches have long been effective at reducing latency and backend load by storing popular content close to users, exploiting the temporal and spatial locality of human-driven access patterns. However, the rise of AI-generated traffic is challenging this assumption. AI agents such as search crawlers and data scrapers issue large volumes of diverse, low-referrer requests with minimal reuse, which degrade cache effectiveness, interfere with human-relevant content, and increase pressure on backend systems. In this paper, we argue that caching infrastructure must evolve to address this shift. We analyze emerging AI traffic patterns and study their impact on caching performance using a CDN prototype based on Wikimedia's architecture. Our results show that even modest amounts of AI traffic lead to significant cache inefficiency. We envision a workload-aware caching paradigm that serves human and AI traffic through differentiated tiers and policies, preserving responsiveness for users while adapting to the diverse access patterns and requirements of AI workloads.
Yazhuo Zhang, Jinqing Cai, Avani Wildani, Ana Klimovic
SoCC1
2025 Unlocking True Elasticity for the Cloud-Native Era with Dandelion
abstract
Elasticity is fundamental to cloud computing. An elastic platform can quickly allocate resources to match the demand of each workload as it arrives, rather than pre-provisioning resources to meet performance objectives. However, even serverless platforms — which boot sandboxes in 10s to 100s of milliseconds — are not sufficiently elastic to avoid pre-provisioning expensive resources. Today's FaaS platforms provision many extra, idle sandboxes in memory to reduce the occurrence of slow, cold starts. Initializing securely isolated sandboxes with a POSIX-like computing environment that today's cloud users expect is slow as it requires booting a guest OS and configuring networking.
Tom Kuchler, Pinghe Li, Yazhuo Zhang, Lazar Cvetkovich, Boris Goranov, Tobias Stocker, Leon Thomm, Simone Kalbermatter, Tim Notter, Andrea Lattuada 0001, Ana Klimovic
SOSP3
2024 SIEVE is Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web Caches
Yazhuo Zhang, Juncheng Yang, Yao Yue, Ymir Vigfusson, K. V. Rashmi
NSDI1
2023 LatenSeer: Causal Modeling of End-to-End Latency Distributions by Harnessing Distributed Tracing
abstract
End-to-end latency estimation in web applications is crucial for system operators to foresee the effects of potential changes, helping ensure system stability, optimize cost, and improve user experience. However, estimating latency in microservices-based architectures is challenging due to the complex interactions between hundreds or thousands of loosely coupled microservices. Current approaches either track only latency-critical paths or require laborious bespoke instrumentation, which is unrealistic for end-to-end latency estimation in complex systems.
Yazhuo Zhang, Rebecca Isaacs, Yao Yue, Juncheng Yang, Lei Zhang 0223, Ymir Vigfusson
SoCC1
2023 FIFO can be Better than LRU: the Power of Lazy Promotion and Quick Demotion
abstract
LRU has been the basis of cache eviction algorithms for decades, with a plethora of innovations on improving LRU's miss ratio and throughput. While it is well-known that FIFO-based eviction algorithms provide significantly better throughput and scalability, they lag behind LRU on miss ratio, thus, cache efficiency.
Juncheng Yang, Ziyue Qiu, Yazhuo Zhang, Yao Yue, K. V. Rashmi
HotOS3
2023 FIFO queues are all you need for cache eviction
abstract
As a cache eviction algorithm, FIFO has a lot of attractive properties, such as simplicity, speed, scalability, and flash-friendliness. The most prominent criticism of FIFO is its low efficiency (high miss ratio).
Juncheng Yang, Yazhuo Zhang, Ziyue Qiu, Yao Yue, K. V. Rashmi
SOSP2
2018 A deep learning model integrating FCNNs and CRFs for brain tumor segmentation
Xiaomei Zhao, Yihong Wu 0002, Guidong Song, Zhenye Li 0003, Yazhuo Zhang, Yong Fan 0001
Medical Image Anal.5