VLDB 2026 Research / reviewers in the wild / expert
Sheng Cheng 0002
dblp:32/2225-2
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-4686-6120ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai |
NSDI | 4 |
| 2026 | From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Sheng Cheng 0002, Jiamin Cao, Shuhong Zhu, Ennan Zhai, Dennis Cai |
SIGCOMM | 3 |
| 2026 | Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001 |
SIGCOMM | 3 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 19 |
| 2024 | Inferring Video Streaming Quality of Real-Time Communication Inside NetworkabstractReal-time video streaming is getting indispensable in people’s daily life, and poses heavy loads and stringent performance requirements on the network. For Internet Service Providers (ISPs), ensuring high-quality real-time video communication is a widely concerned issue. However, inferring the quality of real-time video streaming based on passively-collected network traffic is a great challenge due to limited information in the User Datagram Protocol (UDP) header and the encryption of the application-level protocol. In this paper, we propose IReaV-T to Infer Real-time Video streaming quality with a generalized Transformer, which understands the intrinsic state of the network and predicts the future real-time video quality. By applying novel embedding methods, IReaV-T could make full use of observed traffic features and distinguish different real-time video applications. Extensive comparative experiments demonstrate the effectiveness of IReaV-T, showing that IReaV-T could predict future real-time video quality with mean squared Video Multimethod Assessment Fusion (VMAF) score error less than 6. Yihang Zhang 0007, Sheng Cheng 0002, Zongming Guo, Xinggong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Rebuffering but not Suffering: Exploring Continuous-Time Quantitative QoE by User's Exiting Behaviors
Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
INFOCOM | 1 |
| 2023 | ABRF: Adaptive BitRate-FEC Joint Control for Real-Time Video StreamingabstractAdaptive Forward Error Correction (AFEC) algorithms are proposed to achieve efficient Forward Error Correction (FEC) in Real-Time Communication (RTC). However, current AFEC approaches suffer two key limitations. 1) They do not consider the squeezing effect of redundancy on source bitrate while the squeezing usually happens in real RTC applications with limited bandwidth. 2) They estimate the future packet loss simply, ignoring some critical features of packet loss like randomness and multi-pattern. These drawbacks stop them from providing better Quality of Experience (QoE) in RTC services. We propose ABRF, a general QoE-oriented Adaptive BitRate-FEC joint control algorithm. ABRF makes predictions on the network loss pattern in the coming time and jointly calculates the optimal bitrate-FEC decision based on a QoE model for real-time video streaming. Moreover, ABRF is equipped with a fast adaptation method which helps it generalize across diverse network environments. In terms of Video Multi-method Assessment Fusion (VMAF), experimental results tell that ABRF decreases VMAF degradation caused by packet loss by 68%-95% compared with other AFEC algorithms in real-time video streaming in real-world Internet. Sheng Cheng 0002, Xinggong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | LightFEC: Network Adaptive FEC with a Lightweight Deep-Learning ApproachabstractNowadays, the interest of real-time video streaming reaches a peak. To deal with the problem of packet loss and optimize users' Quality of Experience (QoE), Forward error correction (FEC) has been studied and applied extensively. The performance of FEC depends on whether the future loss pattern is precisely predicted, while the previous researches have not provided a robust packet loss prediction method. In this work, we propose LightFEC to make accurate and fast prediction of packet loss pattern. By applying long short-term memory (LSTM) networks, clustering algorithms and model compression methods, LightFEC is able to accurately predict packet loss in various network conditions without consuming too much time. According to the results of well-designed experiments, we find out that LightFEC outperforms other schemes on prediction accuracy, which improves the packet recovery ratio while keeping the redundancy ratio at a low level. Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
ACM Multimedia | 2 |
| 2020 | DeepRS: Deep-Learning Based Network-Adaptive FEC for Real-Time Video CommunicationsabstractAs real-time multimedia streaming thriving, Forward Error Correction (FEC) methods have been studied and applied extensively these years. Most of researchers paid their attention to the coding algorithms, attempted to balance the trade off between recovery ratio and delay with fewer redundance. However, when packet loss pattern changes dynamically, the redundance waste is too serious to be ignored. In this work, we propose a novel algorithm which adjusts the redundance ratio of FEC encoder according to the prediction of packet loss. Receivers are additionally required to feedback observed packet loss pattern. Streaming sender collects the feedbacked packet loss pattern and predicts the number of packet loss in the incoming short period. As for implementation, we adopt long short-term memory (LSTM) network as our deep learning algorithm, and exquisitely embed it in our adaptive FEC system. With the extensive experiments, our proposed scheme outperforms other FEC methods greatly both in the simulations and evaluations on traces observed from the real world. Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
ISCAS | 1 |