EDBT 2026 Demo / reviewers in the wild / expert
Yachen Wang
dblp:44/6407
· DBLP profile ↗
19ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 16 · 2 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu |
EuroSys | 15 |
| 2026 | SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading
Xingyi Li 0004, Shangguang Wang, Zhehao Lin, Yinben Xia, Qihang Liu, Xiang Li 0067, Zekun He, Yachen Wang, Xianneng Zou |
NSDI | 15 |
| 2026 | Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie |
NSDI | 10 |
| 2026 | Pegasus: A Data Center Network for Bare-Metal AI CloudabstractToday, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications. Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang |
SIGCOMM | 30 |
| 2026 | Accelerating Hardware/Software Combined Traffic Processing With Fast and Efficient Asynchronous Flow OffloadingabstracteHardware/software (hw/sw) combined systems are necessary to meet modern clouds’ requirements for processing huge amounts of network traffic by efficiently offloading large flows to hardware. However, existing hw/sw flow offloading systems typically perform traffic statistics collection and large flow selection within a time window—a time-window-based approach. Their offloading decision of large flows issynchronizedin the unit of a time window, which is mismatched to the asynchronous and dynamic nature of each flow’s sending rate. Additionally, the flow measurement and selection for large flows are decoupled in these solutions, leading to memory and CPU inefficiency. In this paper, we introduce TAO, a novel solution to the hw/sw combined flow offloading problem byasynchronouslyselecting and offloading flows based on flow table entries. TAO can reactfasterto the rapid dynamics of flows by taking actions at each table entry and ismore efficientby coupling flow measurement and selection into the entry. We have implemented a full-fledged TAO prototype based on the P4 switch and DPDK. Testbed results demonstrate that TAO can offload ∼16% more traffic to hardware, outperforming existing solutions by achieving 42× lower memory overhead. Meanwhile, it reduces software CPU utilization by 66.7% and cuts tail forwarding latency by 95.59% compared to state-of-the-art methods. Xijin Yin, Yuanwei Lu, Xin Zhang 0117, Xingtong Lin, Shengli Zheng, Bangwen Deng, Xianneng Zou, Yachen Wang, Guo Chen 0001 |
IEEE Trans. Netw. | 8 |
| 2025 | Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang |
NSDI | 12 |
| 2025 | Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleabstractThe flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers. Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001 |
SIGCOMM | 19 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 13 |
| 2025 | MMCANet A Multimodal and Cross-Attention Network for Cloud Removal and Exploration of Progressive Remote Sensing Images Restoration AlgorithmabstractIn Earth observation, cloud severely affects the interpretation of optical satellites generated high-resolution images. Cloud-free optical images are vital for downstream tasks such as semantic segmentation and object detection. Thus, the elimination of clouds from optical imagery has emerged as a significant topic in remote sensing. Currently, most existing methods are proposed to leverage the texture information from auxiliary synthetic aperture radar (SAR) images to restore cloud-free images via direct channel merging. However, such a unified feature extraction approach often neglects the inherent distribution disparity between SAR and optical images—the result of differing imaging principles-potentially leading to significant feature loss. To this end, we introduce a network by jointing SAR and optical images multimodal and cross-attention network (MMCANet) to effectively extract multiscale contextual features from SAR imagery and integrate them with optical features. Specifically, instead of simple concatenation of the channels of SAR and optical images, we obtain high-dimensional features from them through independent feature extractors. The integration of these features is facilitated by a cross-attention mechanism that provides a more fine-grained amalgamation of information. Meanwhile, an atrous spatial pyramid pooling (ASPP) module is introduced into the integration of high-level features, which captures multiscale contextual information around clouded areas. In addition, we propose four advanced remote sensing image restoration algorithms that approach image restoration as a series of subtasks, gradually eliminating clouds to enhance performance. Comprehensive assessments show that MMCANet performs well on the SEN 12 MS-CR dataset with peak signal-to-noise ratio (PSNR) of 39.8871, structural similarity index (SSIM) of 0.9672, mean absolute error (MAE) of 0.0081, and spectral angle mapper (SAM) of 2.9884. Yejian Zhou, Jiahui Suo, Yachen Wang, Jie Su 0001, Zhen Hong, Rajiv Ranjan 0001, Lizhe Wang 0001, Zhenyu Wen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | AggDeliv: Aggregating Multiple Wireless Links for Efficient Mobile Live Video DeliveryabstractMobile live-streaming applications with stringent latency and bandwidth requirements have gained tremendous attention in recent years. Encountered with bandwidth insufficiency and congestion instability of the wireless uplinks, multi-access networking provides opportunities to achieve fast and robust connectivity. However, the state-of-the-art multi-path transmission solutions are lack of adaptivity to the heterogeneous and dynamic nature of wireless networks. Meanwhile, the indispensable video coding and transformation bring about extra latency and make the video delivery vulnerable to network throughput fluctuation. This paper presents AggDeliv, a framework that provides efficient and robust multi-path transmission for mobile live video delivery. The key idea is to relate multi-path packet scheduling to congestion control optimization over diverse wireless links and adapt it to the mobile video characteristics. This is achieved by probabilistic packet allocation based on links’ congestion windows, wireless-oriented delay and loss aware congestion control, as well as lightweight video frame coding and network-adaptive frame-packet transformation. Real-world evaluations demonstrate that our framework significantly outperforms the state-of-the-art solutions on aggregate goodput and streaming video bitrate. Jinlong E, Lin He 0004, Zongyi Zhao, Yachen Wang, Gonglong Chen |
INFOCOM | 4 |
| 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing ClusterabstractBig data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions. Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 9 |
| 2024 | MegaTE: Extending WAN Traffic Engineering to Millions of Endpoints in Virtualized CloudabstractIn today's virtualized cloud, containers and virtual machines (VMs) are prevailing methods to deploy applications with different tenant requirements. However, these requirements are at odds with the resource allocation capabilities of conventional networking stacks in wide-area networks (WANs). In particular, existing WAN traffic engineering (TE) systems at the granularity of aggregated traffic flows are not designed to cater to each individual flow. In this paper, we advocate for a radical new approach to extend TE systems to involve millions of virtual instance endpoints. We propose and implement a first-of-its-kind system, called MegaTE, to satisfy the needs of each fine-grained traffic flow at the virtual instance level. At the core of the MegaTE system is the paradigm shift from the top-down centralized control to the bottom-up asynchronous query in the TE control loop, combined with eBPF-based segment routing on the data plane and TE optimization contraction on the control plane. We evaluate MegaTE using flow-level simulations with production traffic traces. Our results show that MegaTE supports 20× more endpoints with the similar algorithm run time compared to prior work. MegaTE has been adopted by large-scale public cloud providers. Notably, Tencent rolled out MegaTE in its cloud WAN since December 2022. Our production analysis shows that MegaTE reduces the packet latency of real-time applications by up to 51%. Congcong Miao, Zhizhen Zhong, Yunming Xiao, Senkuo Zhang, Yinan Jiang, Zizhuo Bai, Chaodong Lu, Jingyi Geng, Zekun He, Yachen Wang, Xianneng Zou, Chuanchuan Yang |
SIGCOMM | 11 |
| 2023 | TENSOR: Lightweight BGP Non-Stop RoutingabstractAs the solitary inter-domain protocol, BGP plays an important role in today's Internet. Its failures threaten network stability and will usually result in large-scale packet losses. Thus, the non-stop routing (NSR) capability that protects inter-domain connectivity from being disrupted by various failures, is critical to any Autonomous System (AS) operator. Replicating the BGP and underlying TCP connection status is key to realizing NSR. But existing NSR solutions, which heavily rely on OS kernel modifications, have become impractical due to providers' adoption of virtualized network gateways for better scalability and manageability. Congcong Miao, Yunming Xiao, Marco Canini, Ruiqiang Dai, Shengli Zheng, Jilong Wang 0001, Jiwu Bu, Aleksandar Kuzmanovic, Yachen Wang |
SIGCOMM | 9 |
| 2022 | An Interference-Oriented 5G Radio Resource Allocation Framework for Ultradense NetworksabstractTo cope with the explosive growth in demands of wireless network, ultradense network (UDN) technology is widely adopted, which could increase the capacity of wireless network, but also bring severe intercell interference (ICI). However, existing solutions cannot work well in such complex scenarios, due to the limits of their mechanisms. To solve the problem, in this article, an interference-oriented radio resource allocation framework is proposed with multiple usages, including supplying precise, stable, and timely performance feedbacks, near perfect offline training, and high compatibility. As the use of the framework is derived from precise interference identification, a practical regression-based interference modeling algorithm is proposed to support the framework. With in-depth analysis of the mechanism of interference, the proposed algorithm could efficiently and accurately model interference between users using only data collected from operating wireless networks. Compared with the baseline algorithm, the proposed algorithm could reach the same accuracy with training time of two orders of magnitude shorter. To further show the advantages of the framework, a high-performance double-deep-$Q$-network-based resource allocation algorithm is also proposed. By integrating into the proposed framework, the proposed algorithm could coordinate ICI better, with 40% to 101% higher energy efficiency compared with baseline algorithms. Tao Peng 0001, Yachen Wang, Gonglong Chen |
IEEE Internet Things J. | 3 |
| 2021 | A XGBoost Based Wireless Interference Relation Mining and Performance Prediction MethodabstractUltra-dense network (UDN) is considered to be the key technology for the fifth generation (5G) networks to provide high capacity. However, intensive deployment of femtocells bring severe inter-cells interference (ICI), which greatly limits the performance of the network and the capacity gain the system can obtain. Therefore, the key to solve this problem is to obtain accurate interference information through accurate interference modeling. In fact, the wireless big data generated during the operation of the wireless network contains rich wireless interference information. Based on this, this paper proposes an uplink interference identification and signal-to-interference-plus-noise ratio (SINR) prediction algorithm based on XGBoost and interference model. The proposed algorithm uses the wireless big data generated during network operation to train the XGBoost algorithm, mining the signal-to-interference ratio (SIR) and signal-to-noise ratio (SNR) information between links in the wireless network without increasing the overhead of wireless resources, and then combining with the proposed interference model to achieve accurate prediction of the SINR. The simulation results show that when the training data of the target user reaches 5000 pieces, the prediction error of its SINR will be reduced to less than 0.5dB, which effectively reduces the requirement of data quantity and computing power, and can meet the practical application requirements. Tao Peng 0001, Yachen Wang, Gonglong Chen |
VTC Fall | 4 |
| 2009 | Robust energy detection in cognitive radioabstractThe success of advanced dynamic utilisation of the scarce spectrum in cognitive radio depends upon reliable primary signal detection where accurate noise power estimation plays a critical role. However, in practical scenarios, the noise power cannot be accurately estimated, which significantly degrades the performance of primary signal detection. To avoid inaccurate noise power estimation and associated accumulated problems. A novel two-stage Bayesian estimation-based energy detection algorithm is introduced here. This algorithm, as supported by simulation results, shows two main features: (a) a superior performance of 1 dB compared with previous methods; (b) the consistency of the algorithm has been proved indicating that 100% correct primary user signal detection can be approached as the number of samples tends to infinity. Junyang Shen, Siyang Liu 0001, Yachen Wang, Gang Xie 0002, Habib F. Rashvand |
IET Commun. | 3 |
| 2008 | Full Diversity Spreading code for Downlink Space-Time-Frequency Spreading CDMAabstractRecently proposed downlink space-time-frequency spreading CDMA (STFS-CDMA) is investigated in this paper. By analyzing, a spreading code design criterion of it is derived. From the criterion, we can see that the original two classes of spreading code adopted in STFS-CDMA, Walsh-Hadamard code (WHC) and double-orthogonal code (DOC), both can not always achieve the full space and frequency diversity. Then, a novel spreading code, zero padded rotary FFT code (ZPRFC), is proposed. Compared with WHC coded STFS-CDMA (WHC-STFS-CDMA) and DOC coded STFS-CDMA (DOC-STFS-CDMA), the novel ZPRFC coded STFS-CDMA (ZPRFC-STFS-CDMA) can obtain the full space and frequency diversity at the cost of the reduction of the number of supporting users. Siyang Liu 0001, Yachen Wang |
ICC | 4 |
| 2008 | Optimized Power Control and Resource Allocation in Grouped MC-CDMA SystemsabstractThis paper investigates an optimized dynamic power control for downlink grouped multi carrier-code division multiple access (MC-CDMA) systems to facilitate a high overall system capacity. In the grouped MC-CDMA system, we partition all the subcarriers into several groups with equal space to fully exploit frequency diversity. The optimal power control for maximum system mutual information is introduced based on the optimal subcarrier allocation. However, due to the polynomial problem of transmit power, the accurate solution cannot be obtained. Thankfully, Taylor series expansion enables us to arrive at the approximate optimal power control scheme. According to the analysis result, we propose an approximate optimal power allocation algorithm, which is demonstrated to perform much better than equal power allocation on system capacity. Yachen Wang, Junyang Shen, Siyang Liu 0001 |
ICC | 1 |
| 2008 | Flexible Orthogonal Code for CDMA Based MIMO-OFDM Systems with Space-Time-Frequency SpreadingabstractWe propose a novel spreading code, called flexible-orthogonal code (FOC), for space-time-frequency spreading (STFS) scheme in code division multiple access (CDMA) based multiple input multiple output (MIMO) orthogonal frequency division multiplexing (OFDM). The goal of our design is to obtain an improved bit-error rate (BER) performance without loss of bandwidth efficiency. We make use of multiple antennas to achieve space diversity and balanced signal to interference and noise ratio (SINR) among users on each subscriber device. The proposed code exhibits length flexibility, which breaks up the "power of two" constraint in spreading factor. Our results indicate that the proposed code improves overall performance of the system, by exploiting both space diversity and SINR balancing. The flexibility in code length also results in high spectral efficiency. Yachen Wang, Junyang Shen, Siyang Liu 0001 |
ICC | 1 |