VLDB 2026 Research / reviewers in the wild / expert
Yinben Xia
dblp:155/5692
· DBLP profile ↗
12ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0005-2816-9777ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu |
EuroSys | 11 |
| 2026 | SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading
Xingyi Li 0004, Shangguang Wang, Zhehao Lin, Yinben Xia, Qihang Liu, Xiang Li 0067, Zekun He, Yachen Wang, Xianneng Zou |
NSDI | 8 |
| 2026 | Pegasus: A Data Center Network for Bare-Metal AI CloudabstractToday, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications. Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang |
SIGCOMM | 12 |
| 2025 | Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleabstractThe flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers. Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001 |
SIGCOMM | 16 |
| 2023 | MARB: Bridge the Semantic Gap between Operating System and Application Memory Access BehaviorabstractThe virtual memory subsystem (VMS) is a long-standing and integral part of an operating system (OS). It plays a vital role in enabling remote memory systems over fast data center networks and is promising in terms of transparency and generality. Specifically, these systems use three VMS mechanisms: demand paging, page swapping, and page prefetching. However, the VMS inherent data path is costly, which takes a huge toll on performance. Despite prior efforts to propose page swapping and prefetching algorithms to minimize the occurrences of the data path, they still fall short due to the semantic gap between the OS and applications - the VMS has limited knowledge of its running applications' memory access behaviors. In this paper, orthogonal to prior efforts, we take a fundamen-tally different approach by building an efficient framework to collect full memory access traces at the local bus, and make them available to the OS through CPU cache. Consequently, the page swapping and page prefetching can use this trace to make better decisions, thereby improving the overall performance of systems. We implement a proof-of-concept prototype on commodity x86 servers using a hardware-based memory tracking tool. To show-case our framework's benefits, we integrate it with a state-of-the-art remote memory system and the default kernel page eviction subsystem. Our evaluation shows promising improvements. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yisong Chang, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
DATE | 8 |
| 2023 | HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated MemoryabstractMemory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem’s swapping interface. However, due to the semantic gap between OS and applications – OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults’ access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we buildHoPP– a hardware-software co-designed prefetching framework.HoPPadds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits inHoPP’s software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardware-based memory tracking tool calledHMTTto emulate a modified memory controller. Results show that compared to Fastswap and Leap,HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
HPCA | 7 |
| 2021 | HierCC: Hierarchical RDMA Congestion ControlabstractRDMA has been increasingly deployed in data centers to decrease latency and CPU utilization. However, existing RDMA congestion control schemes fail to address instantaneous large queue build-up or bandwidth under-utilization associated with frequent traffic bursty. In this paper, we argue that traffic uncertainty is the essential reason that constrains data center congestion control from simultaneously achieving high throughput and deterministic latency. Since aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack and hierarchically controls the rate of flows. The rate of aggregate flows between racks is controlled by a credit-based congestion control mechanism. Then the bandwidth obtained by an aggregate flow in a rack is allocated to the corresponding individual flows from that rack promptly and accurately. We evaluate HierCC using SystemC and large-scale NS3 simulations. Results indicate that HierCC can significantly mitigate buffer usage and reduce the 99th percentile FCT by up to 20% and 40% compared with HPCC and DCQCN under a realistic workload, respectively. Jiao Zhang 0002, Zixuan Guan, Zirui Wan, Yinben Xia, Tian Pan 0001, Tao Huang 0005, Dezhi Tang |
APNet | 5 |
| 2021 | ACC: automatic ECN tuning for high-speed datacenter networksabstractFor the widely deployed ECN-based congestion control schemes, the marking threshold is the key to deliver high bandwidth and low latency. However, due to traffic dynamics in the high-speed production networks, it is difficult to maintain persistent performance by using the static ECN setting. To meet the operational challenge, in this paper we report the design and implementation of an automatic run-time optimization scheme, ACC, which leverages the multi-agent reinforcement learning technique to dynamically adjust the marking threshold at each switch. The proposed approach works in a distributed fashion and combines offline and online training to adapt to dynamic traffic patterns. It can be easily deployed based on the common features supported by major commodity switching chips. Both testbed experiments and large-scale simulations have shown that ACC achieves low flow completion time (FCT) for both mice flows and elephant flows at line-rate. Under heterogeneous production environments with 300 machines, compared with the well-tuned static ECN settings, ACC achieves up to 20\% improvement on IOPS and 30\% lower FCT for storage service. ACC has been applied in high-speed datacenter networks and significantly simplifies the network operations. Xiaoliang Wang 0001, Yinben Xia, Derui Liu, Weishan Deng |
SIGCOMM | 4 |
| 2014 | EMD-Based Multi-Model Prediction for Network Traffic in Software-Defined NetworksabstractAccurately predicting for network traffic is significant for network operation and maintenance in software-defined networks (SDN). In this paper, Multi-frequency characteristic of complex network traffic is considered, and a new algorithm named EMD-based multi-model Prediction (EMD-MMP) for network prediction is proposed. The main idea in this algorithm is to decompose the network traffic series into different modes with different frequency by Empirical Mode Decomposition (EMD). According to the characteristics and the cross correlation coefficient of the modes, we reconstruct new components for de-noising by summing up parts of the high frequency modes. Then the new components and the remaining old modes are predicted by ARMA and SVR methods. Finally, the historical traffic data of Internet2 is employed for our experiments to demonstrate the precision of our new prediction algorithm compared with the Auto-Regressive and Moving Average (ARMA) and Support Vector Regression (SVR) models. On average, the EMD-MMP method improves ARMA and SVR by 0.62% and 10.6% at the Mean Absolute Percentage Error (MAPE) statistic indicator, and the Mean Square Error (MSE) of EMD-MMP is 12060.92 while the ARMA and SVR are 13968.8 and 47588.3. Besides, the EMD-MMP algorithm gives a better understanding of the nature of the network traffic. Longfei Dai, Wenguo Yang, Suixiang Gao, Yinben Xia, Mingming Zhu, Zhigang Ji |
MASS | 4 |
| 2014 | OFBGP: A Scalable, Highly Available BGP Architecture for SDNabstractSoftware-Defined Networking (SDN), enabled by OpenFlow, represents a paradigm shift from traditional network to the future Internet. Compared with traditional networks, its controllers and network applications have higher requirements on the availability and scalability. This paper proposes OFBGP, a new scalable, highly available architecture of BGP totally based on distributed design, without any centralized component. OFBGP is an application of SDN controller. In the design of OFBGP, BGP functionalities are divided according to their characteristics and requirements into two modules: BGP Protocol and BGP Decision. By doing so, the application can easily scale out when the number of BGP neighbors increases, and the process of BGP routing decisions can be parallelized. The paper also presents a new solution to implement BGP NSR through snapshots and backup event flow so that from a historical snapshot, replaying real-time backup event flow can rebuild the state that has happened behind at any moment. When a fault occurs, we can easily restore a state to ensure BGP neighbors do not perceive failure. We have also built a prototype to demonstrate our architecture. After the test of the performance of some core modules, the result shows that our new architecture can help improve the availability and scalability of BGP. Wenbo Duan, Deguo Li, Yuanhao Zhou, Yinben Xia, Mingming Zhu |
MASS | 7 |
| 2014 | High availability for Non-stop network controllerabstractNetwork controller is the core of the OpenFlow-Based Software-Defined Networks (SDN). High availability of network controller is an urgent need objectively. The appearance of controller instance fault should not be perceived by data plane. During the fault recovery, OpenFlow messages, especially asynchronous message from data plane, should not be discarded. Otherwise the control platform would hold the outdated network status. This is an interesting problem which we call Asynchronous Message Discard (ASMD) trouble. This paper proposes a Non-stop network controller (NSNC). It can troubleshoot the ASMD problem based on highly available TCP (HA-TCP) technology. During the fault recovery on distributed control platform, we want to elect the optimal controller to take over the network devices managed by the failed controller instance. We analyses the minimum metrics set for controller election. The view of this paper is verified on Floodlight controller. Experiments show that when the controller instance fails, asynchronous messages from switch will not be discarded, and using HA-TCP only adds negligible overhead to throughput and latency. We test the performance of election algorithm based on Mininet. Currently, there is only a theoretical analysis of metrics related to the election. We will test the validity of the election strategy in a real scenario as part of our future work. Deguo Li, Mingfa Zhu, Wenbo Duan, Yuanhao Zhou, Mengxi Chen, Yinben Xia, Mingming Zhu |
WoWMoM | 8 |
| 2014 | Energy-aware routing algorithms in Software-Defined NetworksabstractThe feature of centralized network control logic in Software-Defined Networks (SDNs) paves a way for green energy saving. Router power consumption attracts a wide spread attention in terms of energy saving. Most research on it is at component level or link level, i.e. each router is independent of energy saving. In this paper, we study global power management at network level by rerouting traffic through different paths to adjust the workload of links when the network is relatively idle. We construct the expand network topology according to routers' connection. A 0-1 integer linear programming model is formulated to minimize the power of integrated chassis and linecards that are used while putting idle ones to sleep under constraints of link utilization and packet delay. We proposed two algorithms to solve this problem: alternative greedy algorithm and global greedy algorithm. For comparison, we also use Cplex method just as GreenTe does since GreenTe is state of the art global traffic engineering mechanism. Simulation results on synthetic topologies and real topology of CERNET show the effectiveness of our algorithms. Suixiang Gao, Wenguo Yang, Yinben Xia, Mingming Zhu |
WoWMoM | 5 |