VLDB 2026 Research / reviewers in the wild / expert
Yuchen Xu 0003
dblp:214/3905-3
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-5765-3825ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu |
EuroSys | 1 |
| 2026 | TurboTSS: A Packet Classifier with Fast Rule Lookup and Update for the Cloud
Shaoke Fang, Yuchen Xu 0003, Weize Gao, Jianglong Nie, Wenfei Wu |
INFOCOM | 3 |
| 2026 | Enabling Flexible and Efficient Collective Communication Scheduling in Distributed AI
Yuchen Xu 0003, Xiting Ju, Wenfei Wu |
LANMAN | 1 |
| 2026 | Turbo: Efficiently Serving Long-Context Large Language Models with In-Network AggregationabstractLLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss. Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao |
SIGCOMM | 2 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 7 |
| 2026 | INARouting: Efficient Multi-Job Routing Optimization for Hierarchical In-Network AggregationabstractIn-network aggregation (INA) has emerged as a key technology to alleviate communication bottlenecks in large-scale distributed training, but its performance is often hindered by suboptimal routing. Existing INA-aware routing algorithms suffer from certain limitations: they either lack a global, multi-job coordination mechanism, or operate on incomplete network models that ignore key hardware constraints such as switch processing capacity. These deficiencies lead to network congestion and inefficient resource utilization, ultimately undermining the full potential of INA. To address these challenges, we present INARouting, a novel framework that holistically solves the multi-job hierarchical aggregation routing problem. We propose TINA, a hierarchical aggregation protocol that supports multi-job in-network aggregation. To address different deployment scenarios, we develop two variants: INARouting-Opt that provides optimal solutions for moderate-scale networks, and INARouting-Relax, a fast and effective heuristic using LP-relaxation and a greedy score-based rounding algorithm for large-scale deployments. Through extensive experiments on various scales of Fat-Tree and Spine-Leaf topologies, we demonstrate that INARouting significantly outperforms state-of-the-art methods. INARouting- Opt achieves provably optimal solutions, reducing average job completion time by up to 56% compared to existing methods. Meanwhile, INARouting-Relax outperforms existing algorithms while being 5× faster in solving time, enabling efficient routing in large-scale, dynamic environments. Jianglong Nie, Yidan Yuan, Yuchen Xu 0003, Yitao Yuan, Kehan Yao, Lu Lu 0016, Xiaodong Duan, Wenfei Wu |
IEEE Trans. Netw. | 3 |
| 2025 | MimoSketch: A Framework for Frequency-Based Mining Tasks on Multiple Nodes With SketchesabstractIn distributed data stream mining, we abstract a MIMO scenario where a stream ofmultipleitems is mined bymultiple nodes. We design a framework named MimoSketch for the MIMO-specific scenario, which improves the fundamental mining tasks of item frequency estimation, item size distribution estimation, heavy hitter detection, heavy change detection, and entropy estimation. MimoSketch consists of an algorithm design and a policy to schedule items to nodes. MimoSketch's algorithm applies random counting to preserve a mathematically provenunbiasednessproperty, which makes it friendly to the aggregate query on multiple nodes; its memory layout isdynamicallyadaptive to the runtime item size distribution, which maximizes the estimation accuracy by storing more items. MimoSketch's scheduling policy balances items among nodes, avoiding nodes being overloaded or underloaded, which improves the overall mining accuracy. Our prototype and evaluation show that our algorithm can improve the accuracy of five typical mining tasks by an order of magnitude compared with the state-of-the-art solutions, and the scheduling policy further promotes the performance in MIMO scenarios. Wenfei Wu, Yuchen Xu 0003 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | MimoSketch: A Framework to Mine Item Frequency on Multiple Nodes with SketchesabstractWe abstract a MIMO scenario in distributed data stream mining, where a stream of multiple items is mined by multiple nodes. We design a framework named MimoSketch for the MIMO-specific scenario, which improves the fundamental mining task of item frequency estimation. MimoSketch consists of an algorithm design and a policy to schedule items to nodes. MimoSketch's algorithm applies random counting to preserve a mathematically proven unbiasedness property, which makes it friendly to the aggregate query on multiple nodes; its memory layout is dynamically adaptive to the runtime item size distribution, which maximizes the estimation accuracy by storing more items. MimoSketch's scheduling policy balances items among nodes, avoiding nodes being overloaded or underloaded, which improves the overall mining accuracy. Our prototype and evaluation show that our algorithm can improve the item frequency estimation accuracy by an order of magnitude compared with the state-of-the-art solutions, and the scheduling policy further promotes the performance in MIMO scenarios. Yuchen Xu 0003, Wenfei Wu, Bohan Zhao, Tong Yang 0003, Yikai Zhao 0001 |
KDD | 1 |
| 2023 | Cuckoo Counter: Adaptive Structure of Counters for Accurate Frequency and Top-k EstimationabstractFrequency estimation and top-k flows identification are fundamental problems in network traffic measurement. Sketch, as a basic probabilistic data structure, has been extensively investigated and used in different management applications. However, few of them is suitable for both estimating frequency and finding top-k flows due to the unbalanced distribution of real-world network streams. By introducing a pre-filtering stage to isolate elephant and mice flows, the recently proposed Augmented Sketch (ASketch) significantly improves accuracy for both tasks. However, it suffers from serious performance degradation because of frequent flow exchanges. In this paper, we propose Cuckoo Counter (CC), an adaptive structure that consists of several buckets organized in a specific way. The size of the entry in each bucket is carefully designed to match the actual distribution of streams. During processing, CC hashes a flow to buckets and uses the idea of cuckoo hashing to relocate the flow if an overflow or collision happens, which contributes to fully utilizing memory. Therefore, the replacement strategy helps CC precisely record elephant flows and cover more mice flows, and also guarantees the throughput. Extensive experimental results show that CC has the highest (Freq.) accuracy, excellent (Heavy hitter / change) accuracy, highest (Top-k) precision, and competitive throughput compared to the state-of-the-art. Specifically, CC improves the throughput and accuracy by around 1 and 2 orders of magnitude respectively compared to the well-known ASketch. Qilong Shi, Yuchen Xu 0003, Jiuhua Qi, Wenjun Li 0004, Tong Yang 0003, Yang Xu 0010, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 2 |