Yu Guan 0005

dblp:86/6151-5 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-0726-3933ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai
NSDI1
2025 Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai
NSDI7
2025 Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai
NSDI17
2025 SyCCL: Exploiting Symmetry for Efficient Collective Communication Scheduling
abstract
The performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts.
Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai
SIGCOMM8
2025 3DGS-Enabled High-Fidelity Low-Cost Immersive Static 3D Video Streaming
abstract
3D Gaussian Splatting (3DGS), as the cutting-edge static three-dimensional (3D) content generation technology, revolutionizes the speed and fidelity of 3D model construction and provides immense potential for various applications, including e-commerce, 3D exhibitions, and virtual tourism. However, our pioneering analysis of firsthand user experiments uncovers a critical challenge: the unique user behavior patterns of static 3D scenarios render existing immersive video streaming solutions inadequate. To be concrete, the frequent switches between active and inactive states impair viewport prediction accuracy, while the fast glance and slow view pattern provides an opportunity for further quality of experience (QoE) improvement. To tackle these problems, this paper introduces innovative designs for static 3D video streaming. Specifically, we devise a viewport prediction and error correction mechanism on the client side to restore the user viewport with a low cost. Furthermore, we design a dynamic frame rate and bitrate control algorithm to improve user QoE under various network conditions. We implement the first 3DGS-enabled immersive static 3D video streaming system based on an edge-rendered architecture, ensuring efficient rendering and encoding on the edge server while providing broad accessibility for various client-side devices through a web browser. Extensive testing under real-world network conditions and with various kinds of devices demonstrates that the proposed approach exhibits robust and rapid adaptability to fluctuating network conditions, improving user QoE by over 20%, reducing interactive latency by 89%, and minimizing the stall duration by 26% compared to existing low-latency streaming solutions.
Rongji Liao, Yuan Zhang 0013, Wei Zhang 0324, Lingjun Pu, Yu Guan 0005, Yunpeng Jing, Tao Lin 0001, Jinyao Yan
IEEE J. Sel. Areas Commun.5
2024 Crux: GPU-Efficient Communication Scheduling for Deep Learning Training
abstract
Deep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI.
Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai
SIGCOMM2
2024 Alibaba HPN: A Data Center Network for Large Language Model Training
abstract
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production.
Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai
SIGCOMM6
2023 ZGaming: Zero-Latency 3D Cloud Gaming by Image Prediction
abstract
In cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction.
Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang
SIGCOMM2
2022 Sovereign: Self-Contained Smart Home With Data-Centric Network and Security
abstract
Recent years have witnessed the rapid deployment of smart homes; most of them are controlled by remote servers in the cloud. Such designs raise security and privacy concerns for end users. In this article, we describe the design of Sovereign, a home Internet of Things (IoT) system framework that provides end users complete control of their home IoT systems. Sovereign lets home IoT devices and applications communicate via application-named data and secures data directly. This approach enables direct, secure, one-to-one, and one-to-many Device-to-Device communication over wireless broadcast media. Sovereign utilizes semantic names to construct usable security solutions. We implement Sovereign as a publish–subscribe-based development platform together with a prototype home IoT controller. Our preliminary evaluation shows that Sovereign provides a systematic, easy-to-use solution to user-controlled, self-contained smart homes running on existing IoT hardware without imposing noticeable overhead.
Zhiyi Zhang 0001, Yu Guan 0005, Philipp Moll, Lixia Zhang 0001
IEEE Internet Things J.4
2022 STC: FoV Tracking Enabled High-Quality 16K VR Video Streaming on Mobile Platforms
abstract
The ultra-high-definition 16K Virtual Reality (VR) video is coming to ages with more ”real” virtual experience and less cybersickness. However, the huge bitrate and decoding overhead would overwhelm today’s network and mobile hardware. The widely-known Field-of-View (FoV) adaptation streaming method still has severe bitrate wastes and decoding overhead as it delivers FoV areas with grid-like static tiles. Inspired by this, we present a novel ShiftTile-traCking (STC) streaming scheme, which crops and delivers tiles by tracking FoV movement. It is equivalent to deliver an FoV planar video instead of VR videos. This would save huge bit-rate and reduce decoding complexity. We mainly entail three contributions. 1) To reduce projection distortions, a novel FoV-centric sphere projection is proposed, which projects VR videos with the center of users’ FoVs. 2) To cover diverse FoV movement trajectories with a limited number of tiles, we propose an optimal tiling algorithm by trajectory clustering. 3) To be resilient to FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which scales FoV areas by the prediction accuracy. The evaluation shows that under the same network conditions, STC improves up to 1.3dB V-PSNR, reduces up to 13.2% buffering ratio, and achieves 60% faster decoding speed (61.5 frames per second) compared with state-of-the-art solutions.
Chengyuan Zheng, Jinyu Yin, Fangzhen Wei, Yu Guan 0005, Zongming Guo, Xinggong Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2021 PrefCache: Edge Cache Admission With User Preference Learning for Video Content Distribution
abstract
With the deployment of video streaming in 4G/5G mobile network, Content Delivery Networks (CDN) are extending to the network edge to provide end-users better Quality of Experience (QoE). However, small cache size and irregular request patterns make it a great challenge for edge caching in video content distribution. Most of the existing cache policies are item-wise, they admit each video object separately, which performs poorly on the network edge due to irregular request patterns. We observe that compared with single video objects, users' preferences for video topics are much more constant, thus are easier to be predicted. So we propose PrefCache, a novel cache admission policy based on preference learning, for video content edge caching. PrefCache enables an edge cache to learn users' preferences for videos in real-time. Once receiving a video object, PrefCache decides whether to admit it to the cache by whether it is under users' preference. We make three contributions in this work. (1) First, we design an information collector, which can proactively collect the preference-related information without any modification of clients and video providers. (2) Second, we propose a tree-structure model to learn and compress users' preferences. (3) Third, to decide which videos should be admitted to the cache in real-time, an explore-and-exploit method is applied. We carried out extensive experiments with 24 hours of trace data from a large commercial video content provider. The experimental results demonstrate that PrefCache can improve hit ratio up to 12%, and save 92% memory / 98% CPU overhead, compared to the state-of-the-art cache policies.
Yu Guan 0005, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.1
2020 STC: Enabling 16K VR streaming on mobile platforms with FoV tracking
abstract
16K VR videos are coming to ages. But it could overwhelm mobile hardware for its huge bandwidth consumption and decoding complexity. To enable 16K VR video streaming over mobile platforms, we present a novel ShiftTile-Tracking (STC) streaming system, which crops and transmits video by tracking the Field-of-View (FoV) movement of users. The video chunk is split into ShiftTiles with frame granularity, which always covers FoV areas along the FoV movement trajectory. This transforms a 360-degree VR video into a traditional planar video, which leads to huge bandwidth saving and faster decoding speed. In the system design, we mainly entail two contributions. 1) To accommodate various FoV movement trajectories with a limited number of ShiftTiles, we propose an optimal tiling algorithm by FoV trajectory clustering. 2) To be resilient to the FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which expands the FoV area if the FoV prediction errors are high. The evaluation shows that under the same real-world 4G network conditions, the proposed STC improves 0. 9dBV-PSNR, reduces 12.4% buffering ratio, and achieves 45% faster decoding speed (64 frames per second) on average compared with the state-of-the-art solutions. This enables 16K VR video streaming on current mobile platforms.
Chengyuan Zheng, Jinyu Yin, Yu Guan 0005, Xinggong Zhang, Zongming Guo
GLOBECOM3
2019 UtilCache: Effectively and Practicably Reducing Link Cost in Information-Centric Network
abstract
Minimizing total link cost in Information-Centric Network (ICN) by optimizing content placement is challenging in both effectiveness and practicality. To attain better performance, upstream link cost caused by a cache miss should be considered in addition to content popularity. To make it more practicable, a content placement strategy is supposed to be distributed, adaptive, with low coordination overhead as well as low computational complexity. In this paper, we present such a content placement strategy, UtilCache, that is both effective and practicable. UtilCache is compatible with any cache replacement policy. When the cache replacement policy tends to maintain popular contents, UtilCache attains low link cost. In terms of practicality, UtilCache introduces little coordination overhead because of piggybacked collaborative messages, and its computational complexity depends mainly on content replacement policy, which means it can be O(1) when working with LRU. Evaluations prove the effectiveness of UtilCache, as it saves nearly 40% link cost more than current ICN design.
Lemei Huang, Yu Guan 0005, Xinggong Zhang, Zongming Guo
ICC2
2019 CACA: Learning-based Content-aware Cache Admission for Video Content in Edge Caching
abstract
In the last decades, network caches (Content Distribution Network, CDN) have been widely deployed in video delivery system. As cache has been pushed to network edge as far as possible, small cache size and irregular request pattern make it a great challenge for edge cache to catch popular video contents. Although we can apply cache admission policies to block cold contents out, however, all current admission policies are still based on request pattern (content size, frequency), which perform poorly in edge cache. This paper proposes a novel feature-based cache admission policy, Content-feature Aware Cache Admission(CACA). It admits video objects to cache by video features, not by request pattern anymore. The intuition behind that is, for a group of users, their preferred contents may change at any time, but their preferred content features would maintain for a while. Popularity of video features (such as topic, author), is much more predicable than that of single video object. To mine critical features from huge feature space, this paper proposes a tree-structure reinforcement learning algorithm. Critical features are learned from a feature-partition tree which is spanned and pruned by history popularity. Then, an Exploration-and-Exploitation method is used to select the Top-K critical features. Video contents with these features will be admitted to cache. We carried out extensive experiments with 24-hours data traces from a commercial video content provider. The experimental results demonstrate that the proposed CACA is able to improve hit ratio up to 15%, reduce back-to-origin up to 20% and save 95% memory, compared with state-of-art cache admission policies.
Yu Guan 0005, Xinggong Zhang, Zongming Guo
ACM Multimedia1
2019 Pano: optimizing 360° video streaming with a better understanding of quality perception
abstract
Streaming 360° videos requires more bandwidth than non-360° videos. This is because current solutions assume that users perceive the quality of 360° videos in the same way they perceive the quality of non-360° videos. This means the bandwidth demand must be proportional to the size of the user's field of view. However, we found several quality-determining factors unique to 360° videos, which can help reduce the bandwidth demand. They include the moving speed of a user's viewpoint (center of the user's field of view), the recent change of video luminance, and the difference in depth-of-fields of visual objects around the viewpoint.
Yu Guan 0005, Chengyuan Zheng, Xinggong Zhang, Zongming Guo, Junchen Jiang
SIGCOMM1
2018 Name-Based Routing with On-Path Name Lookup in Information-Centric Network
abstract
Name-based routing is one of the core ideas in Information-centric network (ICN). In name-based routing, there is a tradeoff between the cost of name announcement and name lookup. Some ICN architectures introduce an efficient way of name lookup but pay high price in name announcement, others cut off most information exchange in name announcement yet introduce heavy burden in name lookup. In order to solve this problem and balance the cost of name announcement and lookup, we propose Name-based routing with On-Path Name Lookup (OPNL). OPNL looks up name prefixes on the path to name's guaranteed destination. It accomplishes distributed name lookup with lighter burden while maintaining little information exchange in name announcement. Results of simulation experiments show that OPNL makes a tradeoff between the cost of name announcement and lookup to have better scalability, eliminates storage overhead and communication overhead compared with prior works and attains even better performance.
Yu Guan 0005, Lemei Huang, Xinggong Zhang, Zongming Guo
ICC1
2017 A Caching Miss Ratio Aware Path Selection Algorithm for Information-Centric Networks
abstract
In Information-Centric Networks (ICN), contents are cached on some intermediary routers. This creates thus a new situation which is totally different from the traditional path-selection paradigm: the source/destination paradigm no longer exists; instead, the new paradigm is how to find a path through a selected group of caches, so that the content is delivered via the shortest way. This paper addresses this issue and proposes a path-selection algorithm taking into account both the caching capability of router and the more traditional link cost between routers. We formulated the problem as a convex optimization problem (named ESP) which aims to get expected shortest path (ESP) by minimizing the transportation cost. By applying the Lagrangian dual theorem, we solved the ESP problem and obtained a criterion for request (and reversely, data) routing. Based on this path-selection criterion, we provide a fully distributed distance-based ESP algorithm that enables routers maintain routes to nearest content, without knowing a network topology and the caching miss ratio of content at other routers. Simulations confirm the efficiency of our approach versus the traditional shortest path algorithm.
Weihong Lin, Xinggong Zhang, Yu Guan 0005, Zongming Guo
LCN3