EDBT 2026 Demo / reviewers in the wild / expert
Zhiyi Yao
dblp:180/6996
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM Training
Leyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu 0001, Zeng Chuxuan, Ling Deng, Hui Wang 0011 |
INFOCOM | 2 |
| 2026 | Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie |
NSDI | 2 |
| 2026 | Modeling Point-to-Point Dependency for High-Dimensional Long-Term Series Forecasting
Xinyu Li 0014, Kexi Chen, Ying Zheng 0004, Zhiyi Yao, Yi Xie 0003, Jihan Dai, Lei Bai 0001, Jin Zhao 0001, Jiajie Shen, Yunqi Cai, Hong Lu 0001, Xin Wang 0002 |
WWW | 4 |
| 2026 | An Efficient Computing and Communication Framework for Large-Scale Data Processing Cluster
Xuya Jia, Zhiyi Yao, Edison Liu, Congcong Miao, Yuedong Xu 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | MemFerry: A Fast and Memory Efficient Offload Training Framework with Hybrid GPU Computation
Zhiyi Yao, Zuning Liang, Yuedong Xu 0001, Jin Zhao 0001, Hui Wang 0011 |
INFOCOM | 1 |
| 2025 | Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang |
NSDI | 1 |
| 2025 | Minerva: Decentralized Collaborative Query Processing Over InterPlanetary File SystemabstractData silos create barriers to accessing and utilizing data dispersed over networks. Directly sharing data easily suffers from the long downloading time, the single point failure and the untraceable data usage. In this paper, we present Minerva, a peer-to-peer cross-cluster data query system based on the InterPlanetary File System (IPFS). Minerva makes use of the distributed Hash table (DHT) lookup to pinpoint the locations that store content chunks. We theoretically model the DHT query delay and introduce a fat Merkle tree structure as well as the DHT caching to reduce it. We design the query plan for read and write operations on top of Apache Drill that enables the collaborative query with decentralized workers. We conduct comprehensive experiments on Minerva, and the results show that Minerva achieves up to$2.08 \times$query performance acceleration compared to the original IPFS data query, and can complete data analysis queries on the Internet-like environments within an average latency of 0.615 second. With a collaborative query, Minerva could perform up to$1.39 \times$performance acceleration than the centralized query with raw data shipment. Zhiyi Yao, Qianlan Bai, Yuedong Xu 0001 |
IEEE Trans. Big Data | 1 |
| 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing ClusterabstractBig data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions. Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 2 |
| 2021 | Vehicle-to-Vehicle Channel Characterization Based on Ray-Tracing for Urban Road ScenariosabstractIn this paper, the vehicle‐to‐vehicle (V2V) channel characteristics in peak hours at the 5.9 GHz band in two typical urban road scenarios, the urban straight road and the intersection, are investigated. The channel characteristics, such as path loss, root mean square (RMS) delay spread, and angular spread, are derived from the ray‐tracing (RT) simulations. Due to the low height of antennas at both the transmitter (Tx) and the receiver (Rx), the line of sight (LOS) between the Tx and the Rx will often be obstructed by other vehicles. Based on the RT simulation results, the shadowing loss is modelled by the multimodal Gaussian distribution, and path loss models in both LOS and non‐LOS (NLOS) conditions are obtained. And the RMS delay spread in two scenarios can be modelled by the Weibull distribution. In addition, the deployment of an antenna array is discussed based on the statistics distribution of the angular spread. Zhiyi Yao, Haiyang Miao, Bo Ai 0001 |
Wirel. Commun. Mob. Comput. | 2 |